University · Buying · Lesson 4
How to read a Spark accelerator benchmark
ETL Compare staff · Published 19 September 2026 · 2 min read
In brief
A speed or cost figure means little without its baseline, workload, data size, cluster and software versions. Check each of those, note who ran the test, and convert the figure to cost per run on one of your own jobs before comparing products.
What was the baseline?
The same product can look very different against different starting points. Some figures are against open-source Apache Spark. AWS states that the Amazon EMR runtime for Spark is up to 5.4x faster than open-source Spark, so a team already on EMR starts from a different place. DualBird's Iceberg compaction benchmark names two baselines, vanilla Spark and "state-of-the-art C++ accelerated Spark". Always find the baseline before reading the multiplier.
What workload was measured?
Most open-source projects publish query suites derived from TPC-H and TPC-DS, which stress scans, joins and aggregations. Comet's benchmarking guide says it runs "benchmarks derived from TPC-H and TPC-DS", and Gluten's site reports TPC-H and TPC-DS results. These are useful for comparing engines, but an ETL job often spends more of its time in shuffle, spill and writes. A workload shaped like ETL, such as table compaction, tests those stages more directly.
What size, cluster and versions?
Look for the data scale (for example 1 TB or 3 TB), node count and instance type, Spark version and accelerator version. Results age quickly: Gluten's published Velox backend figures were tested in June 2023, and in June 2026 two AWS engineers reported on the DataFusion blog that Comet went from 11% slower than vanilla Spark to 32% faster on their 3 TB TPC-DS run after two months of fixes with the maintainers.
Who ran it, and can you reproduce it?
Vendor and project figures are a starting point, not a measurement of your job. Prefer results that publish scripts and configuration, results from someone other than the vendor, and results at more than one cluster size. Then take the figure to your own pipeline: our proof of concept guide and cost per job show how.
Related
Sources
- Amazon EMR · Fetched 27 Sep 2026
- DualBird Iceberg compaction benchmark · Fetched 27 Sep 2026
- Comet benchmarking guide · Fetched 27 Sep 2026
- Apache Gluten · Fetched 27 Sep 2026
- Apache Gluten on GitHub · Fetched 27 Sep 2026
- Comet on Amazon EKS · Fetched 27 Sep 2026