Guide · Figures checked September 2026
How to profile a slow Spark ETL job before buying anything
ETL Compare staff · Figures checked September 2026
In brief
Before you buy an accelerator, find the stage that is slow. Open the Spark UI, find the stage that takes the most time, and check four things in its task metrics: shuffle read and fetch wait (shuffle), Shuffle spill (disk) (spill), the gap between median and max task duration (skew), and the number and size of input and output files (small files). Each points to a different fix, and only some are fixed by an accelerator.
| Accelerator | Read | Transform | Shuffle | Spill | Write |
|---|---|---|---|---|---|
| Gluten + Velox | Partial | Documented | Documented | Partial | Not stated |
| Comet | Documented | Documented | Documented | Partial | Not stated |
| RAPIDS Accelerator | Documented | Documented | Documented | Not stated | Documented |
| DualBird | Partial | Documented | Documented | Documented | Partial |
| Auron | Partial | Documented | Documented | Partial | Not stated |
| Photon | Documented | Documented | Not stated | Not stated | Documented |
| Flarion | Partial | Documented | Not stated | Not stated | Not stated |
- Documented: the vendor's own documentation says it addresses this stage
- Partial: indirect or experimental, or covered only by an end-to-end claim
- Not stated: not found on the pages we reviewed
This shows what each vendor says, not what we measured. Sources are listed on each review.
Step 1. Which stage is actually slow?
Open the Spark UI for one representative run of the job (the Spark History Server for a finished run). In the Jobs tab, find the job that takes most of the wall-clock time, then open its stages. Sort the stages by duration. In most ETL jobs one or two stages account for most of the runtime; everything after this step is about those stages.
Record four numbers per slow stage for your baseline: duration, number of tasks, input size, and shuffle read and write size. You will compare against them after every change, including any accelerator trial.
Step 2. Is the job waiting on shuffle?
Open the slow stage and look at the task metrics the Spark UI reports: Shuffle Read Size / Records, Shuffle Read Fetch Wait Time, Shuffle Remote Reads and Shuffle Write Time.
| Signal | What it suggests |
|---|---|
| High Shuffle Read Fetch Wait Time relative to task duration | Tasks are blocked waiting for shuffle data from other executors: network or remote disk is the bottleneck, not compute. |
| Shuffle write size close to or larger than the stage input | A wide transformation (join, groupBy, distinct, repartition) is moving most of the data; check whether a broadcast join or an earlier filter would avoid it. |
| Very many tiny shuffle partitions | The partition count is too high for the data size. |
Configs to check
spark.sql.shuffle.partitions (default 200): the number of partitions used when shuffling for joins or aggregations.spark.sql.adaptive.enabled (default true) and spark.sql.adaptive.coalescePartitions.enabled (default true): adaptive query execution merges small shuffle partitions at runtime.spark.sql.autoBroadcastJoinThreshold (default 10 MB): tables below this size are broadcast instead of shuffled.
What an accelerator can do: several products in this guide document shuffle work (Gluten's ColumnarShuffleManager, Comet's native shuffle, the RAPIDS Shuffle Manager, Auron's compacted shuffle format, and DualBird, which states it reduces shuffle data). None of them removes a shuffle that the query plan requires. If an avoidable join is shuffling, fix the plan first.
Sources: spark.apache.org/docs/latest/web-ui.html, spark.apache.org/docs/latest/sql-performance-tuning.html · Fetched 27 Sep 2026
Step 3. Is the job spilling to disk?
In the same stage view, check Shuffle spill (memory) and Shuffle spill (disk). The Spark UI defines Shuffle spill (memory) as the size of the deserialized form of the shuffled data in memory, and Shuffle spill (disk) as the size of the serialized form of the data on disk. Any non-zero disk spill on a slow stage means tasks ran out of execution memory and wrote intermediate data to local disk, then read it back.
Check Peak execution memory and GC time next. High GC time alongside spill points to memory pressure on the executors.
Why spill keeps coming back after you add memory: executors share execution memory among the tasks running on them, so adding cores per executor can reduce memory per task even when total memory grows. DualBird's article 'A Simpler Spark' walks through this sharing rule with a 4-task, 16 GiB example.
Common fixes without new software: fewer concurrent tasks per executor (fewer cores per executor or more memory per core), more shuffle partitions so each task handles less data, and removing skew (step 4). What an accelerator can do: DualBird states it eliminates disk spills; Gluten and Comet document native spill with limits (experimental in Gluten's Velox backend, bounded per task in Comet); Auron documents multi-level memory management. Check these claims on your worst-spilling stage.
Sources: spark.apache.org/docs/latest/web-ui.html, dualbird.io/resources/blog/a-simpler-spark · Fetched 27 Sep 2026
Step 4. Is the data skewed?
Open the summary metrics for the slow stage and compare the median and max task duration, and the median and max shuffle read size. If the max is many times the median, a few partitions hold far more data than the rest, and the stage waits on those few tasks (stragglers).
Adaptive query execution handles one common case. With spark.sql.adaptive.skewJoin.enabled (default true), Spark splits skewed partitions in sort-merge joins. A partition counts as skewed when it is larger than spark.sql.adaptive.skewJoin.skewedPartitionFactor (default 5.0) times the median partition size and also larger than spark.sql.adaptive.skewJoin.skewedPartitionThresholdInBytes (default 256MB). Skew in aggregations or in joins below those thresholds is not covered.
Other fixes: salting hot keys, filtering or pre-aggregating hot keys separately, and broadcasting the small side of a join. Skew and spill often come together: the oversized task is also the one that spills. An accelerator does not change the distribution of your keys; the question to ask a vendor is what happens to the oversized task.
Sources: spark.apache.org/docs/latest/sql-performance-tuning.html · Fetched 27 Sep 2026
Step 5. Are small files slowing reads and writes?
Look at the first (read) stage: if the number of tasks is very high and the input per task is tiny, the job is spending its time opening files and planning, not processing. spark.sql.files.maxPartitionBytes (default 128 MB) sets the maximum bytes packed into one read partition for file sources such as Parquet, JSON and ORC.
On the write side, count the files a job produces per partition. Thousands of small output files slow every downstream reader. Fixes: coalesce or repartition before writing, and for table formats such as Apache Iceberg, run compaction. Compaction is itself a heavy Spark job; DualBird's public benchmark happens to be a sort-based Iceberg compaction job, which makes it a relevant workload to ask vendors about.
Step 6. What should you ask an accelerator vendor, and how should you test?
- Which of my slow stages does your product address, and what falls back to standard Spark?
- Does it need code changes, a platform move or a different instance type?
- What is the cost per job on my workload, including your fee and any change in instance price?
- Can we test on one of my real pipelines, at two or three cluster sizes, against my current runtime (for example the Amazon EMR runtime, not open-source Spark)?
- Who supports a failed job at 3 a.m., and what happens if we remove the plugin?
Run the trial against the baseline you recorded in step 1. Compare stage by stage: if the accelerator speeds up a stage that was not your bottleneck, the total will barely move.
Frequently asked questions
What is the fastest way to see why a Spark job is slow?
Open the Spark UI, sort the stages by duration, and in the slowest stage check shuffle fetch wait time, Shuffle spill (disk), and the gap between median and max task duration.
Does adding executor memory stop disk spill?
Not always. Execution memory is shared by the tasks running on an executor, so more cores per executor can mean less memory per task. Check spill per task, not only total memory.
Does adaptive query execution fix skew?
Partly. It splits skewed partitions in sort-merge joins when they exceed both the skew factor (default 5.0 times the median) and the size threshold (default 256MB). Skewed aggregations are not covered.
Should I profile before trying an accelerator?
Yes. Without a stage-level baseline you cannot tell whether an accelerator fixed your bottleneck or sped up a stage that did not matter.
Sources
- Web UI - Spark documentation · Fetched 27 Sep 2026
- Performance Tuning - Spark SQL documentation · Fetched 27 Sep 2026
- Faster Spark Was the Goal. A Simpler Spark Came With It · Fetched 27 Sep 2026