ETL CompareETL Compare

University · Basics · Lesson 1

How Spark runs a job: jobs, stages and tasks

ETL Compare staff · Published 21 September 2026 · 2 min read

In brief

Spark turns each action in your code into a job, splits the job into stages wherever data has to be shuffled, and runs each stage as tasks, one per partition. Stage boundaries are where shuffle and spill costs appear, which is why both profiling and accelerator claims are best read stage by stage.

What starts a job?

Spark evaluates transformations lazily. Nothing runs until an action asks for a result, such as writing a table, counting rows or collecting data to the driver. Each action becomes a job. An ETL pipeline that writes three output tables therefore runs at least three jobs, and the Spark UI lists them separately in the Jobs tab.

Where does one stage end and the next begin?

Narrow transformations such as filter, select and simple column expressions work on each partition on its own, so Spark chains them inside a single stage. Wide transformations such as joins, aggregations, distinct and repartition need rows with the same key in the same partition. Spark ends the stage there, writes the shuffle output, and starts a new stage that fetches it. That boundary is the shuffle.

A broadcast join is the main exception: when one side is below spark.sql.autoBroadcastJoinThreshold (default 10 MB), Spark sends it to every executor and avoids shuffling the large side.

How many tasks does a stage run?

One per partition. On the read stage, file splits decide the count; spark.sql.files.maxPartitionBytes (default 128 MB) caps how much data goes into one partition. After a shuffle, spark.sql.shuffle.partitions (default 200) sets the count, and adaptive query execution can merge small partitions at runtime. Tasks run inside executors, and the tasks on one executor share its execution memory, so a partition that is much larger than the rest can spill to disk or become a straggler.

Why does this matter when you compare accelerators?

The Spark UI Stages tab reports shuffle read and write, spill, GC time and task duration per stage, so your profile of a slow job is a list of stages. Accelerator vendors also describe their products by stage: some document faster scans and operators, some target shuffle, some target spill. Matching the two lists is the fastest way to see whether a product addresses your bottleneck. The profiling guide walks through it.

Related

Sources