ETL CompareETL Compare

Notes · Published 27 September 2026

How to run a Spark accelerator proof of concept on your own ETL pipeline

ETL Compare staff · Published 27 September 2026 · Figures checked September 2026

  • How-to
  • Buying

A trial plan that produces a decision, not a demo: one real pipeline, a logged baseline, three cluster sizes and a threshold set in advance.

In brief

Pick one real pipeline, record a baseline from Spark event logs on the platform you run today, then run the accelerator on the same data at two or three cluster sizes and compare stage by stage and cost per run. Set the pass or fail threshold before the first run, check how much of the plan fell back to standard Spark, and plan time for one tuning round, because first results often move.

Which pipeline should you test?

Choose a job that is expensive, runs on a schedule and represents your real work, and whose slow stages you have already found with the profiling guide. A job that spends its time in shuffle and spill tests a different claim from one that spends it in scans. Use real data if you can: accelerator coverage depends on the operators, data types and file formats your job actually uses, and a synthetic job can miss the parts that fall back.

Write down what better means before you start. Lower cost per run, a shorter runtime to meet a deadline, or both, lead to different cluster sizes and different winners.

How do you record a baseline?

Turn on event logging for every baseline and trial run. The Spark monitoring documentation says that setting spark.eventLog.enabled to true "configures Spark to log Spark events that encode the information displayed in the UI to persisted storage". The History Server rebuilds the Spark UI from those logs, by default on port 18080, and the same metrics are available as JSON from the REST API under /api/v1, including per-stage task summaries.

  • Wall-clock runtime and cluster up-time you pay for
  • Node count, instance type and hourly price
  • Per-stage duration, shuffle read and write, and Shuffle spill (disk)
  • Median and max task duration on the slowest stage
  • Input and output file counts and bytes

Run the baseline at least three times on the same input and use the median. Measure on the runtime you use today. AWS states that the Amazon EMR runtime for Spark is up to 5.4x faster than open-source Apache Spark, so a trial measured against open-source Spark tells an EMR team little about its own saving.

What should the test matrix look like?

Run the accelerator on the same input at two or three cluster sizes, not one. A single point can hide whether a product scales or only wins at one configuration. Keep the Spark version, input data and output target identical to the baseline. Where a product needs a different instance type, choose node counts by hourly cluster cost rather than by node count, so each trial run has a clear cost comparison.

An example test matrix (illustrative; set your own sizes)
RunClusterWhat it tests
Baseline, three runsYour current clusterReference runtime, stage profile and cost per run
Trial AAbout half the current hourly cluster costWhether the product saves money at a smaller footprint
Trial BAbout the same hourly cluster costThe cleanest like-for-like comparison
Trial CAbout 1.5 times the current hourly cluster costWhether a faster runtime is worth paying for when there is a deadline

How do you check what actually ran accelerated?

Every product we track falls back to standard Spark for work it cannot run natively. The RAPIDS Accelerator documents automatic CPU fallback, Photon falls back to the Spark runtime, and Gluten and Comet document automatic fallback. A job can finish with the accelerator switched on and still run most of its heavy stages in the JVM. Ask the vendor or project how to see which operators ran natively in the physical plan or the Spark UI SQL tab, and record the share for your slowest stages. The RAPIDS Accelerator also ships a Qualification Tool that estimates fit from your existing event logs before you run anything.

Check correctness as well as speed. Compare row counts and checksums of the output tables between baseline and trial runs. In an August 2026 post on Spark upgrades, Flarion argues that proving output data is unchanged, not changing code, is the hard part of a migration; the same discipline applies when you add an execution engine. See fallback.

How do you turn the runs into a cost decision?

Use cost per run, not speedup. Cost per run is the cluster's total hourly cost (instances, platform fee and any accelerator fee) times the runtime in hours, plus storage and transfer, as set out in Spark cost per job, explained. The break-even speedup is the new hourly cluster cost divided by the old one. Ask commercial vendors for a written quote before the trial ends: DualBird and Flarion do not publish prices on the pages we reviewed, and the open-source plugins cost nothing to license but need engineering time to configure.

Why might the first result not be the final one?

Early numbers move. In a June 2026 post on the Apache DataFusion blog, two AWS engineers describe running a 3 TB TPC-DS workload on Spark 3.5.8 on Amazon EKS. They report that their first Comet run was 11% slower than vanilla Spark and that, after two months of work with the maintainers, the final run was 32% faster overall. Plan the trial with time for one tuning round, keep every run's settings in the log, and give the baseline the same tuning attention so the comparison stays fair.

What is a fair pass or fail rule?

  1. Set the threshold in cost per run and in runtime before the first run.
  2. Pass only if the median of three runs clears it at a cluster size you would actually buy.
  3. Confirm the output tables match the baseline.
  4. Confirm the slowest stages ran natively, not in fallback.
  5. Record the support path: who answers when a job fails, and how you remove the plugin.

Sources

Related