ETL CompareETL Compare

University · Basics · Lesson 2

How Spark accelerators plug into Spark

ETL Compare staff · Published 20 September 2026 · 3 min read

In brief

Most Spark accelerators do not change your code. They load through Spark configuration, rewrite the physical plan so that supported operators run on a native engine, often replace the shuffle manager, and hand anything unsupported back to standard Spark. The differences between products are in the engine underneath, the memory it needs and how much of your plan falls back.

How does an accelerator attach to a Spark job?

Through configuration and a jar on the classpath. Gluten is enabled with spark.plugins=org.apache.gluten.GlutenPlugin, off-heap memory switched on and a columnar shuffle manager. Comet is added from Maven Central with its extension, off-heap memory and CometShuffleManager. Auron is switched on with spark.auron.enable and a custom shuffle manager. The RAPIDS Accelerator is a plugin that, in its FAQ's words, replaces parts of the physical plan that Apache Spark considers internal. On Databricks, Photon is a runtime setting (a checkbox or runtime_engine=PHOTON on classic compute, and the default on serverless). DualBird's product page describes two steps: change the instance type and use its Spark plugin. Flarion states it is added through configuration parameters.

What runs underneath?

A native engine outside the JVM. Gluten passes plans to Velox, a C++ engine created at Meta, or to ClickHouse. Comet and Auron run on Apache DataFusion, a Rust engine built on Apache Arrow, and Flarion describes a DataFusion-based, Arrow-native engine. Photon is a native C++ engine inside the Databricks Runtime. The RAPIDS Accelerator runs supported operations through NVIDIA's RAPIDS libraries. See vectorized execution and columnar execution for why native, batch-oriented engines are faster at the same work.

What happens to operations the engine cannot run?

They fall back to standard Spark. Every product in this guide documents some form of fallback, so a job keeps working even when coverage is partial. The cost is hidden: data may have to be converted between columnar and row formats at each boundary, and a heavy stage that falls back gets no benefit. When you trial a product, check which operators in your slowest stages actually ran natively.

Why does memory configuration change?

Native engines usually work in off-heap memory, which you size yourself with spark.memory.offHeap.size. Gluten's getting-started example uses 20g. Comet's tuning guide notes that its memory accounting "isn't 100% accurate" and describes the pool fractions and spill limits to set. Replacing the shuffle manager also changes how shuffle data is written, which is why several products document shuffle as a separate feature.

Related

Sources