University · Buying · Lesson 5
How to match a Spark accelerator to your slowest stage
ETL Compare staff · Published 17 September 2026 · 3 min read
In brief
Profile the job first, name the part of the job and the metric that dominate its runtime, then check which accelerators document addressing that part. A product that speeds up operators will not fix a job that waits on shuffle fetches, and a product aimed at spill matters less on a job that never spills.
Which five parts of a job do we track?
This site splits an ETL job into five parts: read (scan and decode files), transform (filter, join, aggregate), shuffle (write and fetch data between stages), spill (memory pressure that pushes data to disk) and write (encode and commit output). They are not Spark's own stage boundaries, which fall at shuffles, but they match how vendors describe what they speed up and how the Spark UI reports time.
How do you tell which part dominates?
| Part of the job | What to look at | What a problem looks like |
|---|---|---|
| Read | Input size and duration of the first stages; input file and task counts | Many small input files, or long scan stages on large tables |
| Transform | Task duration and GC time on stages without large shuffles | Uniformly slow tasks with little shuffle or spill |
| Shuffle | Shuffle Write Size and Time, Shuffle Read Size, Shuffle Read Fetch Wait Time, Shuffle Remote Reads | High fetch wait time: tasks sitting idle waiting for data |
| Spill | Shuffle spill (memory), Shuffle spill (disk), Peak execution memory | Any disk spill on a slow stage, often concentrated in a few large tasks |
| Write | Duration of the final stage, output file count and size | Long final stages or many small output files |
The profiling guide walks through each metric step by step.
Which accelerators document which part?
The grid below comes from the same stage data used across the site. "Documented" means the vendor's or project's own documentation says the product addresses that part of the job; it is a statement about documentation, not a measurement.
| Accelerator | Read | Transform | Shuffle | Spill | Write |
|---|---|---|---|---|---|
| Gluten + Velox | Partial | Documented | Documented | Partial | Not stated |
| Comet | Documented | Documented | Documented | Partial | Not stated |
| RAPIDS Accelerator | Documented | Documented | Documented | Not stated | Documented |
| DualBird | Partial | Documented | Documented | Documented | Partial |
| Auron | Partial | Documented | Documented | Partial | Not stated |
| Photon | Documented | Documented | Not stated | Not stated | Documented |
| Flarion | Partial | Documented | Not stated | Not stated | Not stated |
- Documented: the vendor's own documentation says it addresses this stage
- Partial: indirect or experimental, or covered only by an end-to-end claim
- Not stated: not found on the pages we reviewed
This shows what each vendor says, not what we measured. Sources are listed on each review.
Coverage differs most outside the transform column, so start with the column that matches your slowest stage, then read the stage note on each product's review before you shortlist.
What if your bottleneck is not covered?
Ask the vendor or project directly, and ask for evidence: a benchmark on a workload of the same shape, or a trial on your job. "Not stated" in our grid means we did not find it on the pages we reviewed, not that the product cannot help. Also check the fixes that need no product: broadcast joins for small tables, fewer and larger files, repartitioning skewed keys, and the adaptive query execution settings.
Why does the bottleneck change the business case?
An accelerator's saving applies to the part of the runtime it addresses. If shuffle and spill take half of a job's runtime and a product only speeds up transform work, the best case is bounded by the other half. That is why this site weights stage coverage at 20% in its method, and why the proof of concept guide compares runs stage by stage, not only end to end.
Related
Sources
- Web UI - Spark documentation · Fetched 27 Sep 2026
- Performance Tuning - Spark SQL documentation · Fetched 27 Sep 2026