Sail

Last updated: Mon Oct 5 15:20:46 2026

Warning

Sail support is experimental. It is only available in the development version of pysparklyr.

pak::pak("mlverse/pysparklyr")

Intro

Sail is a Spark Connect server written in Rust. It does not need Java or a JVM. Sail runs the same Spark DataFrame and SQL commands that Spark does, so sparklyr can talk to it the same way it talks to Spark Connect. You can connect to a remote Sail server that is already running, or to a local Sail server on your machine.

Get started

To start and work with a local Sail server, use "sail" as the method. If you can use uv on your machine, sparklyr installs the Python libraries that Sail needs when you connect. It then starts a local Sail server and connects you to it. You do not need to do any other setup (a.k.a. no Java needed!).

library(sparklyr)

sc <- spark_connect(master = "local", method = "sail")
#> Retrieving version from PyPi.org
#> ✔ PyPi specs: 'pyspark-client' version 4.2.0, requires Python >=3.10 [50ms]
#> 
#> ℹ
#> ✔ Python environment: 'Managed `uv` environment' [943ms]
#> 
#> [2026-10-05T20:20:48Z INFO sail_python::spark::server] Starting the Spark Connect server on 127.0.0.1:57003...
#> [2026-10-05T20:20:48Z INFO sail_session::session_manager::actor::handler] creating session 1fa24ec4-596b-4746-826e-7013e9e2fa3b

The server uses a free port on your machine. spark_disconnect() stops the server. The server also stops when the R session ends.

Python environment

If you cannot use uv, run install_sail() once, before you connect. It creates a Python environment that stays on your machine. These are the main libraries in the new Python environment:

  • pyspark-client, which is a thin version of pyspark
  • pysail, which sparklyr uses to start local Sail servers
  • rpy2, which runs the R code that spark_apply() sends
pysparklyr::install_sail()

After that, connect with spark_connect() as shown in Get started. sparklyr finds and uses the new environment, instead of having uv create a temporary one.

To install a specific Sail version, pass it in version:

pysparklyr::install_sail(version = "0.7")

Work with data

Sail works with the same sparklyr and dplyr code that you use with Spark. Copy a local data frame to Sail with copy_to():

library(dplyr)

mtcars_tbl <- copy_to(sc, mtcars, memory = FALSE)

Then use dplyr verbs. sparklyr turns them into Spark SQL, and Sail runs the query:

mtcars_tbl |>
  group_by(am) |>
  summarise(mpg = mean(mpg, na.rm = TRUE))
#> # A query:  ?? x 2
#> # Database: connect_sail
#>      am   mpg
#>   <dbl> <dbl>
#> 1     1  24.4
#> 2     0  17.1

Use collect() to bring the results into R:

mtcars_tbl |>
  filter(hp > 150) |>
  select(mpg, cyl, hp) |>
  collect()
#> # A tibble: 13 × 3
#>      mpg   cyl    hp
#>    <dbl> <dbl> <dbl>
#>  1  18.7     8   175
#>  2  14.3     8   245
#>  3  16.4     8   180
#>  4  17.3     8   180
#>  5  15.2     8   180
#>  6  10.4     8   205
#>  7  10.4     8   215
#>  8  14.7     8   230
#>  9  13.3     8   245
#> 10  19.2     8   175
#> 11  15.8     8   264
#> 12  19.7     6   175
#> 13  15       8   335

You can also run SQL directly with DBI:

DBI::dbGetQuery(sc, "SELECT cyl, COUNT(*) AS n FROM mtcars GROUP BY cyl")
#>   cyl  n
#> 1   6  7
#> 2   4 11
#> 3   8 14

Run R code

spark_apply() runs an R function on each group of the data. Here, nrow() counts the rows for each value of am. columns sets the names and types of the result:

mtcars_tbl |>
  spark_apply(nrow, group_by = "am", columns = "am double, x long")
#> # A query:  ?? x 2
#> # Database: connect_sail
#>      am     x
#>   <dbl> <dbl>
#> 1     1    13
#> 2     0    19
Warning

With a "local" connection, the R code runs inside your R session, not in a separate R process. This means:

  • If the R code crashes, it can end your R session
  • The R code can change objects in your global environment

Connect to a running Sail server

You can also connect to a Sail server that is already running. Pass the server’s address in master. The address uses the “sc://” protocol. Your machine still needs the Python environment from Get started or Python environment:

sc <- spark_connect(
  master = "sc://localhost:50051",
  method = "sail"
)

The Python environment that runs the Sail server needs pyspark-client. To use spark_apply(), the server also needs:

  • rpy2
  • R
  • The same Python version that your R session uses

See the Sail documentation to learn how to start a Sail server.

How it works

sparklyr uses reticulate to call the Python pyspark-client library. That library sends the commands to the Sail server over gRPC. With a "local" connection, pysail runs the Sail server, which is written in Rust, inside your R session (Figure 1). With a remote connection, the commands go to a Sail server on another machine instead.

flowchart LR
  subgraph rs[R session]
    subgraph r[R]
      sr[sparklyr]
      rt[reticulate]
    end
    subgraph ps[Python]
      pc[pyspark-client]
      subgraph rust[Rust]
        sl[Sail server]
      end
    end
  end
  sr <--> rt
  rt <--> pc
  pc <-- gRPC --> sl

  style rs  fill:#fff,stroke:#666,color:#000
  style r   fill:#fff,stroke:#666,color:#000
  style sr  fill:#fff,stroke:#666,color:#000
  style rt  fill:#fff,stroke:#666,color:#000
  style ps  fill:#fff,stroke:#666,color:#000
  style pc  fill:#fff,stroke:#666,color:#000
  style rust fill:#fff,stroke:#666,color:#000
  style sl  fill:#fff,stroke:#666,color:#000

Figure 1: How sparklyr communicates with a local Sail server

Limitations

Sail does not support some sparklyr features:

  • Caching - copy_to() and the spark_read_*() functions need memory = FALSE. With memory = TRUE, they return an error. compute() also returns an error.
  • Machine learning - The ml_*() and ft_*() functions do not work.
  • Tuning - tune_grid_spark() does not work.