Glossary
What is a Spark disk spill?
In brief
A disk spill happens when a Spark task needs more execution memory than it has, so it writes intermediate data (for example sort buffers or aggregation state) to local disk and reads it back later. Spill shows up in the Spark UI as Shuffle spill (memory) and Shuffle spill (disk); any disk spill on a slow stage is time spent on disk I/O and serialization instead of computing.
Why does spill happen even on large executors?
Execution memory on an executor is shared by the tasks running on it. As DualBird's article 'A Simpler Spark' explains, with N active tasks each task gets roughly 1/(2N) of the pool before spilling and is capped near 1/N; its example, 4 concurrent tasks sharing 16 GiB, leaves each task about 1 to 2 GiB. A skewed partition larger than that spills, however large the executor.
How do you reduce spill?
Give each task more memory (fewer cores per executor or more memory per core), split the work into more, smaller partitions, fix skew so no single task is oversized, and prefer operations that do not buffer whole partitions.
Which accelerators address spill?
Documented
- DualBird: its product page states it eliminates disk spills
Partial
- Gluten + Velox: experimental spill in the Velox backend
- Comet: native operators spill within a per-task limit
- Auron: multi-level memory management
Not stated
- RAPIDS Accelerator: not stated on the pages we reviewed
- Photon: not stated on the pages we reviewed
- Flarion: not stated on the pages we reviewed
Related terms
Sources
- Web UI - Spark documentation · Fetched 27 Sep 2026
- Faster Spark Was the Goal. A Simpler Spark Came With It · Fetched 27 Sep 2026
- Apache Gluten: Velox backend getting started · Fetched 27 Sep 2026
- Apache DataFusion Comet tuning guide · Fetched 27 Sep 2026
- DualBird product page · Fetched 27 Sep 2026
- Apache Auron on GitHub · Fetched 27 Sep 2026