ETL CompareETL Compare

Notes · Published 26 September 2026

Six common claims about Spark accelerators, checked against public documentation

ETL Compare staff · Published 26 September 2026 · Figures checked September 2026

  • Buying
  • Benchmarks

What vendor and project pages actually say about code changes, platforms, hardware, coverage, benchmarks and cost.

In brief

Most common beliefs about Spark accelerators are half right. None of the seven products we track needs code changes, you do not have to leave Amazon EMR to use one, and only one of them needs GPUs. But no product runs every operator natively, no vendor multiplier predicts your saving, and no license fee does not mean no cost.

Do Spark accelerators need code changes?

The claim: adopting an accelerator means rewriting pipelines.

What the documentation says: none of the seven products we track requires changes to Spark SQL or DataFrame code. Gluten, Comet and Auron load through Spark configuration and a jar. The RAPIDS Accelerator replaces "parts of the physical plan that Apache Spark considers internal". Photon is a runtime setting on Databricks. DualBird states setup is an instance-type change plus its Spark plugin, and Flarion states "zero code and infrastructure changes". Setup still takes work, mostly memory and shuffle configuration for the open-source plugins.

Verdict: true for code; configuration still changes.

Do you have to leave Amazon EMR to accelerate Spark?

The claim: faster Spark means moving off Amazon EMR to another platform.

What the documentation says: no. DualBird, the RAPIDS Accelerator and Flarion list Amazon EMR on the pages we reviewed, and DualBird publishes a guide to setting it up on EMR in two steps. Open-source plugins such as Gluten, Comet and Auron do not publish EMR guides but can generally be added where you control Spark configuration. The one product that implies a platform move is Photon, which only runs on Databricks. EMR's own runtime matters too: AWS states it is up to 5.4x faster than open-source Apache Spark, so measure any accelerator against the EMR runtime you run today.

Verdict: false. EMR is the platform layer, and accelerators run on it.

Do Spark accelerators need GPUs?

The claim: acceleration means GPUs.

What the documentation says: only the RAPIDS Accelerator runs on NVIDIA GPU instances. Gluten, Comet and Auron run on standard CPU instances, Flarion lists standard instances, Photon runs on supported Databricks instance types including Graviton, and DualBird runs on Amazon EC2 F2 instances.

Verdict: false for six of the seven products.

Does an accelerator speed up every part of a job?

The claim: switch it on and the whole job runs faster.

What the documentation says: every product documents fallback to standard Spark for work it cannot run natively, and coverage differs by stage. Databricks documents that Photon does not support UDFs, RDD APIs or stateful streaming. Gluten documents spill in its Velox backend as experimental. Comet's write path is not described on the pages we reviewed. A job can finish with the accelerator switched on and still run its heaviest stage in the JVM. The stage coverage grid shows what each product documents.

Verdict: false. Check which operators ran natively on your own job.

Does a vendor's speedup figure predict your saving?

The claim: a 3x or 10x headline figure is what you will see.

What the documentation says: each figure comes from a different workload, baseline and cluster. Gluten reports TPC-H and TPC-DS results against vanilla Spark. Auron states about 2x faster than Spark 3.5 on TPC-DS. DualBird's compaction benchmark compares vanilla Spark and "state-of-the-art C++ accelerated Spark". Flarion's 3x figure comes without a published methodology on the pages we reviewed. The AWS engineers' Comet report on the DataFusion blog shows how far results can move: 11% slower than vanilla Spark on the first run and 32% faster after two months of fixes, as the authors state.

Verdict: false. Use the figure to decide what to test, not what to expect.

Is an open-source accelerator free?

The claim: Apache-licensed means no cost.

What the documentation says: Gluten, Comet, Auron and the RAPIDS Accelerator have no license fee, but the RAPIDS Accelerator needs GPU instances, and the others need off-heap memory sized and shuffle settings managed by your team, with community support. Comet's tuning guide notes that its memory accounting "isn't 100% accurate". The cost moves from a license into instances and engineering time.

Verdict: half true. Free to license, not free to run.

How should you use these checks?

Treat each claim as a question for the trial rather than a fact to accept. The buyer's checklist turns them into questions to answer before the first run, and the proof of concept guide shows how to test them on your own pipeline.

Sources

Related