ETL CompareETL Compare

Notes · Published 24 September 2026

Apache DataFusion under Spark: how Comet, Auron and Flarion share one engine

ETL Compare staff · Published 24 September 2026 · Figures checked September 2026

  • Engines
  • Open source

Three of the seven accelerators we track run on the same Rust engine. What that means for features, releases and buyers.

In brief

Apache DataFusion is an extensible query engine written in Rust that uses Apache Arrow as its in-memory format. Three Spark accelerators build on it: Apache DataFusion Comet, Apache Auron (formerly Blaze) and the commercial Flarion. Work that lands in DataFusion, such as the pluggable spill backends in 55.0.0, can reach all three, but each product decides how much of Spark it maps onto the engine and how it handles shuffle and memory.

What is Apache DataFusion?

DataFusion describes itself as an "extensible query engine written in Rust that uses Apache Arrow as its in-memory format". Apache Arrow is a columnar memory layout, so DataFusion is a columnar, vectorized engine: its operators work on batches of column values rather than one row at a time. It is a library for building data systems, not a Spark replacement, which is why Spark accelerators can embed it.

How do Comet, Auron and Flarion use it?

DataFusion-based Spark accelerators, as documented on the pages we reviewed
ProductHow it uses DataFusionLicenseStatus
Apache DataFusion CometA Spark plugin that translates Spark physical plans into DataFusion plans for supported operators, with native or columnar shuffleApache License 2.0Subproject of Apache DataFusion; 1.0.0 released 7 August 2026
Apache Auron (formerly Blaze)Maps Spark physical plans onto DataFusion native execution, with its own compacted shuffle format and multi-level memory managementApache License 2.0Incubating at the Apache Software Foundation; latest listed release v8.0.0-incubating
FlarionDescribes a DataFusion-based, Arrow-native execution engine for Spark, Hadoop and RayCommercialCompany; price not on its homepage

The shared engine does not make the products the same. Comet documents its own memory pool, tuning settings and shuffle manager; Auron documents its own shuffle format and memory management; Flarion describes caching, performance monitoring and anomaly detection. Our Comet review, Auron review and Flarion review score each on its own documentation, and DataFusion Comet vs Apache Auron puts two of them side by side.

What did DataFusion 55.0.0 add?

Apache DataFusion 55.0.0 was released on 25 August 2026, from 877 commits by 175 contributors. It adds range partitioning, MERGE INTO planner support, runtime row-group pruning and pluggable spill backends. Row-group pruning lets a scan skip blocks of Parquet data that cannot match a query, and spill backends decide how the engine writes data to disk under memory pressure, which matters for spill-heavy ETL.

An engine release does not reach Spark users until the accelerator built on it adopts it. Comet 1.0.0 (7 August 2026) came out before DataFusion 55.0.0, so check the release notes of Comet, Auron or Flarion for the DataFusion version they ship before counting on a new engine feature.

How does DataFusion compare with Velox?

Both are native, vectorized execution engines used under Spark. Velox is written in C++, was created at Meta and reaches Spark through Apache Gluten, which can also use ClickHouse as a backend. DataFusion is written in Rust on Apache Arrow and reaches Spark through Comet, Auron and Flarion. Both communities were active in the period we track: on 25 August 2026, the day DataFusion 55.0.0 shipped, the Velox blog described collecting Delta Lake statistics natively for Gluten.

What does a shared engine mean for buyers?

  • Engine features are necessary, not sufficient. A DataFusion feature helps only once the accelerator you run ships it and maps your operators onto it.
  • Compare the Spark layer, not only the engine. Plan translation, fallback, shuffle and memory handling are where the three products differ, and they decide how much of your job runs natively.
  • Watch both release calendars. For Comet and Auron, the Apache DataFusion and Apache Spark release notes both affect what you can run.
  • Test on your own job. Engine-level benchmarks say little about a spill-heavy or shuffle-heavy ETL pipeline; our proof of concept guide shows how to test.

Sources

Related