ETL CompareETL Compare

Guide · Figures checked September 2026

Spark cost per job, explained

ETL Compare staff · Figures checked September 2026

In brief

The cost of one Spark ETL run is the hourly price of every node in the cluster times the hours the job runs, plus any platform fee, accelerator fee and storage or transfer charges. An accelerator saves money only if its speedup is larger than the increase in hourly cost it causes. If it needs instances that cost twice as much per hour, it has to run the job more than twice as fast to save anything.

What goes into the cost of one job run?

cost per run = (nodes x instance price per hour + platform fee per hour + accelerator fee per hour) x runtime in hours + storage and data transfer for the run

  • Instance price: what your cloud charges per node-hour for the instance type, on-demand, spot or with commitments.
  • Platform fee: for example the Amazon EMR charge on top of EC2, or Databricks DBUs. Check your platform's pricing page.
  • Accelerator fee: zero for the open-source plugins (Gluten, Comet, Auron, the RAPIDS plugin); a license or usage fee for commercial products, which DualBird and Flarion do not publish on the pages we reviewed.
  • Runtime: the wall-clock hours the cluster is up for this job, including start-up if you pay for it.
  • Storage and transfer: shuffle and spill volumes need local or attached disk; cross-zone transfer can add cost.

What is the break-even speedup?

break-even speedup = new hourly cost of the cluster / old hourly cost of the cluster

If the new setup (instances plus platform fee plus accelerator fee) costs 1.5 times as much per hour, the job must run more than 1.5 times faster to cost less. Anything above that is saving; anything below is a loss, however impressive the speedup sounds.

Worked example

Illustrative numbers for the arithmetic only. They are not prices from any vendor or cloud price list.

Cost per run under six illustrative scenarios
ScenarioHourly cost of clusterRuntimeCost per runVersus baseline
1.Baseline: 10 nodes at $1.00 per node-hour$10.003.0 h$30.00reference
2.Open-source plugin, same instances, 2x faster$10.001.5 h$15.0050% lower
3.Different instance type at $2.00 per node-hour, 10 nodes, 4x faster, no fee$20.000.75 h$15.0050% lower
4.Same as 3 but only 1.8x faster$20.001.67 h$33.3311.1% higher
5.Same as 3 plus a $5.00 per hour accelerator fee, 4x faster$25.000.75 h$18.7537.5% lower
6.Fewer nodes: 5 nodes at $2.00 per node-hour, 3x faster than baseline$10.001.0 h$10.0066.7% lower

Scenario 4 is the one to watch: a real speedup that still loses money because the hourly cost doubled. Scenario 6 is the one vendors will propose: a faster setup often needs fewer nodes, so compare at the node count you would actually run.

How do vendors state cost savings, and how should you check them?

Vendors report savings in different units. DualBird states 50%-90% lower costs on its homepage and reports cost per TB in its Iceberg compaction benchmark. Flarion states a 60% cost reduction. Auron states about 50% cluster resources saved on TPC-DS. Open-source projects mostly report speedups, not cost. None of these are measurements of your job.

Convert every claim to cost per run on your own pipeline: take your baseline cost per run from the formula above, run the trial, and recompute with the trial's node count, instance price, fees and runtime. If the vendor reports cost per TB, divide your baseline cost per run by the TB your job processes to compare like with like.

DualBird states

“10x-30x faster data processing performance with 50%-90% lower costs”

Source: DualBird homepage · Fetched 27 Sep 2026

Flarion states

“60% cost reduction”

Source: Flarion homepage · Fetched 27 Sep 2026

Auron states

“~50% cluster resources” saved

Source: Apache Auron · Fetched 27 Sep 2026

What about the platform baseline?

Measure against what you run today. AWS states that Amazon EMR's runtime is up to 5.4x faster than open-source Apache Spark. A saving measured against open-source Spark can look much smaller against the EMR runtime, and the reverse is also true. Ask every vendor which baseline they used.

Sources: aws.amazon.com/emr · Fetched 27 Sep 2026

Frequently asked questions

How do I calculate the cost of a Spark job?

Multiply the cluster's total hourly cost (instances, platform fee and any accelerator fee) by the job's runtime in hours, then add storage and data transfer for that run.

Does a faster Spark job always cost less?

No. If the faster setup costs more per hour, the speedup has to exceed the increase in hourly cost. A job that runs 1.8x faster on instances that cost 2x as much per hour costs more.

Are open-source Spark accelerators free?

The software is free under the Apache License 2.0. You still pay for compute and for the engineering time to configure and tune it.

Sources

Related