ETL CompareETL Compare

Notes · Published 29 September 2026

A buyer's checklist for Spark ETL accelerators

ETL Compare staff · Published 29 September 2026 · Figures checked September 2026

  • Buying
  • How-to

Twenty questions, in five groups, to answer before the first trial run.

In brief

Before you trial a Spark accelerator, answer five groups of questions: what is slow and why, whether the product runs where your jobs run, what it costs, who supports it, and what evidence stands behind its claims. Most answers are on public pages or in your own Spark UI, and the ones that are not become your questions for the vendor or project.

What is slow, and why?

  1. Which pipelines cost the most per month, and which miss their deadline?
  2. In the Spark UI, which stages dominate those jobs, and is the time in scans, operators, shuffle, spill or writes?
  3. Is the problem skew, small files or a missing broadcast join that a configuration change would fix?
  4. Is adaptive query execution on, and are shuffle partitions sized for the data?

If the last two questions turn up a fix, apply it first and measure again. The profiling guide and the lesson on adaptive query execution cover the usual ones.

Will it run where your jobs run?

  1. Is the product documented for your platform, whether Amazon EMR, Databricks, Google Cloud Dataproc, AWS Glue or self-managed Spark?
  2. Does it need different instances, such as NVIDIA GPU instances for the RAPIDS Accelerator or Amazon EC2 F2 instances for DualBird, and can you get that capacity in your region?
  3. Does it support your exact Spark patch version and JDK?
  4. Does it support your table and file formats, such as Apache Iceberg, Delta Lake and Parquet?

Our platform matrix shows which products list which platforms, and none of the pages we reviewed list AWS Glue. You do not have to leave Amazon EMR to use an accelerator: several list it, and open-source plugins can usually be installed where you control Spark configuration.

What will it cost?

  1. Is there a license fee, and is it published?
  2. Does the hourly instance price change, and by how much?
  3. What speedup would you need to break even, given the new hourly cluster cost divided by the old one?
  4. How much engineering time will tuning and upgrades take?
License and published price terms, as published on the pages we reviewed, 27 September 2026
AcceleratorLicensePublished priceInstances
Gluten + VeloxApache License 2.0No license feeStandard CPU instances (x86_64 or aarch64)
CometApache License 2.0No license feeCommodity CPU instances (amd64 or arm64)
RAPIDS AcceleratorApache License 2.0No plugin fee; GPU instance pricing appliesNVIDIA GPU instances (Volta or later)
DualBirdCommercialNot published; contact DualBirdAmazon EC2 F2 instances
AuronApache License 2.0No license feeStandard CPU instances
PhotonCommercial (part of Databricks)Databricks DBU pricing; Photon changes DBU consumptionSupported Databricks instance types, including Graviton
FlarionCommercialNot on homepage; listed on AWS MarketplaceStandard instances

Sources: gluten.apache.org, Velox backend getting started, apache/incubator-gluten on GitHub, datafusion.apache.org/comet, Comet installation guide, Comet tuning guide, spark-rapids overview, RAPIDS Accelerator user guide, RAPIDS Accelerator FAQ, NVIDIA/spark-rapids on GitHub, dualbird.io, DualBird product page, Iceberg compaction benchmark, A Simpler Spark, auron.apache.org, apache/auron on GitHub, kwai/blaze on GitHub, Photon documentation, flarion.io · Fetched 27 Sep 2026

Work the numbers per run, not per hour: see Spark cost per job, explained and our note on what Spark accelerators cost.

Who supports it?

  1. Is support from a vendor under contract or from community channels?
  2. Who answers when a production job fails at night?
  3. How often does the product release, and how quickly does it follow new Spark releases?
  4. How do you switch it off, and have you tested that?

The lesson on open-source or commercial accelerators compares support models, and running an accelerator in production covers rollback.

What evidence stands behind the claims?

  1. What baseline does the figure use: open-source Spark, a platform runtime such as Amazon EMR's, or another accelerator?
  2. What workload was measured, and does it look like your ETL?
  3. Are cluster size, data size and software versions published?
  4. Has anyone other than the vendor or project published results?

Every product we track publishes some figure, and none can be compared directly with another. Our lesson on reading benchmarks explains why, and the published evidence score on each review summarises what we found for this edition.

What do you do with the answers?

Shortlist the two or three products that fit your platform and bottleneck, then run the same proof of concept on each. The calculator can help with the shortlist: move its sliders to match your answers, for example raising cost model transparency if a public price is a procurement requirement, and see how the ranking changes. The side-by-side table puts every score and reason for your shortlist on one page.

Sources

Related