ETL CompareETL Compare

University · Buying · Lesson 5

How to match a Spark accelerator to your slowest stage

ETL Compare staff · Published 17 September 2026 · 3 min read

In brief

Profile the job first, name the part of the job and the metric that dominate its runtime, then check which accelerators document addressing that part. A product that speeds up operators will not fix a job that waits on shuffle fetches, and a product aimed at spill matters less on a job that never spills.

Which five parts of a job do we track?

This site splits an ETL job into five parts: read (scan and decode files), transform (filter, join, aggregate), shuffle (write and fetch data between stages), spill (memory pressure that pushes data to disk) and write (encode and commit output). They are not Spark's own stage boundaries, which fall at shuffles, but they match how vendors describe what they speed up and how the Spark UI reports time.

How do you tell which part dominates?

Where each part of the job shows up in the Spark UI
Part of the jobWhat to look atWhat a problem looks like
ReadInput size and duration of the first stages; input file and task countsMany small input files, or long scan stages on large tables
TransformTask duration and GC time on stages without large shufflesUniformly slow tasks with little shuffle or spill
ShuffleShuffle Write Size and Time, Shuffle Read Size, Shuffle Read Fetch Wait Time, Shuffle Remote ReadsHigh fetch wait time: tasks sitting idle waiting for data
SpillShuffle spill (memory), Shuffle spill (disk), Peak execution memoryAny disk spill on a slow stage, often concentrated in a few large tasks
WriteDuration of the final stage, output file count and sizeLong final stages or many small output files

The profiling guide walks through each metric step by step.

Which accelerators document which part?

The grid below comes from the same stage data used across the site. "Documented" means the vendor's or project's own documentation says the product addresses that part of the job; it is a statement about documentation, not a measurement.

Spark ETL job stages and which stages each accelerator documents addressingA waterfall of five Spark ETL job stages (read, transform, shuffle, spill, write) with illustrative proportions, and below it a grid showing, for each accelerator, whether its own documentation says it addresses that stage.Anatomy of a Spark ETL jobRead: scan and decode filesTransform: filter, join, aggregateShuffle: write and fetch between stagesSpill: memory pressure pushes data to diskWrite: encode and commit outputIllustrative proportions, not measured data. Your own split comes from the Spark UI: see Profile a slow Spark job.
Stage coverage by accelerator, from vendor documentation
AcceleratorReadTransformShuffleSpillWrite
Gluten + VeloxPartialPartialDocumentedDocumentedDocumentedDocumentedPartialPartialNot statedNot stated
CometDocumentedDocumentedDocumentedDocumentedDocumentedDocumentedPartialPartialNot statedNot stated
RAPIDS AcceleratorDocumentedDocumentedDocumentedDocumentedDocumentedDocumentedNot statedNot statedDocumentedDocumented
DualBirdPartialPartialDocumentedDocumentedDocumentedDocumentedDocumentedDocumentedPartialPartial
AuronPartialPartialDocumentedDocumentedDocumentedDocumentedPartialPartialNot statedNot stated
PhotonDocumentedDocumentedDocumentedDocumentedNot statedNot statedNot statedNot statedDocumentedDocumented
FlarionPartialPartialDocumentedDocumentedNot statedNot statedNot statedNot statedNot statedNot stated
  • DocumentedDocumented: the vendor's own documentation says it addresses this stage
  • PartialPartial: indirect or experimental, or covered only by an end-to-end claim
  • Not statedNot stated: not found on the pages we reviewed

This shows what each vendor says, not what we measured. Sources are listed on each review.

Coverage differs most outside the transform column, so start with the column that matches your slowest stage, then read the stage note on each product's review before you shortlist.

What if your bottleneck is not covered?

Ask the vendor or project directly, and ask for evidence: a benchmark on a workload of the same shape, or a trial on your job. "Not stated" in our grid means we did not find it on the pages we reviewed, not that the product cannot help. Also check the fixes that need no product: broadcast joins for small tables, fewer and larger files, repartitioning skewed keys, and the adaptive query execution settings.

Why does the bottleneck change the business case?

An accelerator's saving applies to the part of the runtime it addresses. If shuffle and spill take half of a job's runtime and a product only speeds up transform work, the best case is bounded by the other half. That is why this site weights stage coverage at 20% in its method, and why the proof of concept guide compares runs stage by stage, not only end to end.

Related

Sources