University · Figures checked September 2026
Spark ETL university: glossary and lessons
In brief
ETL Compare University collects the background a buyer needs before comparing Spark accelerators: nine lessons in three tracks (Basics, Buying and Running it) and a glossary of 38 Spark performance terms. If Spark internals are new to you, start with Basics and read in order, then use the glossary as a reference.
Which track should you start with?
Basics
How Spark runs a job, how accelerators attach to it and how memory is shared between tasks. Start here if Spark internals are new to you.
Lesson 1 · 2 min read
How Spark runs a job: jobs, stages and tasks
Spark turns each action in your code into a job, splits the job into stages wherever data has to be shuffled, and runs each stage as tasks, one per partition.
Lesson 2 · 3 min read
How Spark accelerators plug into Spark
Most Spark accelerators do not change your code.
Lesson 3 · 4 min read
How Spark memory works: executors, tasks and off-heap memory
Each Spark executor has a pool of execution memory that its running tasks share, so the memory one task gets depends on how many tasks run beside it.
Buying
How to read vendor benchmarks, match a product to your slowest stage and weigh open-source against commercial options.
Lesson 4 · 2 min read
How to read a Spark accelerator benchmark
A speed or cost figure means little without its baseline, workload, data size, cluster and software versions.
Lesson 5 · 3 min read
How to match a Spark accelerator to your slowest stage
Profile the job first, name the part of the job and the metric that dominate its runtime, then check which accelerators document addressing that part.
Lesson 6 · 3 min read
Open-source or commercial Spark accelerators: what actually changes
Gluten, Comet, Auron and the RAPIDS Accelerator are released under the Apache License 2.0 with no license fee; DualBird and Flarion are commercial, and Photon is part of Databricks.
Running it
The settings, records and routines that keep an accelerator trial honest and a production deployment stable.
Lesson 7 · 3 min read
Adaptive query execution and the Spark settings that shape shuffle
Adaptive query execution (AQE) re-plans a Spark SQL query while it runs, using statistics from finished stages to merge small shuffle partitions, split skewed ones and switch join strategies.
Lesson 8 · 3 min read
Spark event logs and the History Server: keeping evidence from every run
Turn on Spark event logging for every run you want to compare.
Lesson 9 · 3 min read
Running a Spark accelerator in production: versions, fallback checks and rollback
Once an accelerator is in production, the work shifts from proving a speedup to keeping it: match every Spark, JDK and platform upgrade to the versions the product supports, watch for operators that start falling back, validate output after each change, and keep a tested way to switch the product off.
Where is the glossary?
The glossary defines 38 terms, from adaptive query execution to Velox, with longer pages for shuffle, disk spill, vectorized execution, columnar execution and Velox.