ETL CompareETL Compare

Review · Figures checked September 2026

Apache DataFusion Comet review: a Spark accelerator on DataFusion

ETL Compare staff · Figures checked September 2026 · Sourced from vendor docs, project pages and public benchmarks

In brief

DataFusion Comet scores 3.8 out of 5, second in this edition. It runs Spark operators on the Rust-based Apache DataFusion engine with no code changes, supports Spark 3.5, 4.0 and 4.1, ships prebuilt Linux jars and publishes a TPC-DS benchmark. Its tuning guide is frank that memory accounting is not exact, so plan time for memory settings.

3.8 / 5

ETL fit score (editorial assessment, 0-5)

ETL fit score (editorial assessment, 0-5)

Rank 2 of 7

Best for teams on recent Spark 4.x releases

Comet at a glance
ApproachOpen-source Spark plugin that runs supported operators on the Apache DataFusion engine (Rust, Arrow)
Runs onSelf-managed Spark 3.5, 4.0 and 4.1 (3.4 deprecated) on Linux
InstancesCommodity CPU instances (amd64 or arm64)
LicenseApache License 2.0
PriceNo license fee
Code changesNone; jar plus Spark configuration
Figures checkedSeptember 2026

Sources: datafusion.apache.org/comet, Comet installation guide, Comet tuning guide · Fetched 27 Sep 2026

How does it score on each criterion?

Score breakdown by criterion, with reasons
CriterionWeightScoreWhy
Adoption effort20%3.8No code changes; add the Comet jar from Maven Central, enable the extension and configure off-heap memory and the Comet shuffle manager.
Stage coverage20%3.6Native Parquet scans (including for Iceberg), native operators and native or columnar shuffle; native operators can spill, bounded per task; write path not described on the pages we reviewed.
Tuning and operating burden15%2.8The tuning guide says memory accounting "isn't 100% accurate" and describes pool fractions, batch size and spill limits to set; community support.
Platform and instance portability10%4.2Supports Spark 3.5, 4.0 and 4.1 (3.4 deprecated) with prebuilt Linux amd64 and arm64 jars on commodity hardware; you install it yourself on any platform.
Cost model transparency10%5.0Tied highestApache License 2.0, no license fee.
Published evidence15%3.8Tied highestPublishes a TPC-DS at 1 TB benchmark with a per-query breakdown in its Benchmarking Guide.
Maturity and community10%3.9A subproject of Apache DataFusion that has reached a 1.x release line; younger than Gluten and RAPIDS.

What is DataFusion Comet?

Comet is 'a high-performance accelerator for Apache Spark' that runs existing Spark queries on Apache DataFusion, an extensible query engine written in Rust that uses Apache Arrow as its in-memory format. It is a subproject of Apache DataFusion, Apache 2.0 licensed, and runs on commodity hardware. When a query uses a feature Comet does not support, it falls back to Spark automatically.

Sources: datafusion.apache.org/comet, datafusion.apache.org · Fetched 27 Sep 2026

Which Spark versions and platforms does it support?

The installation guide lists Spark 3.4.3, 3.5.9, 4.0.4 and 4.1.3, with experimental support for 4.2.0; Spark 3.4 support is deprecated as of the 1.0.0 release. Comet needs JDK 17 or later. The jars on Maven Central bundle native libraries for Linux amd64 and arm64.

Sources: Comet installation guide · Fetched 27 Sep 2026

How much tuning does it need?

Comet shares Spark's off-heap memory pool (spark.memory.offHeap.size). The tuning guide says 'Comet's memory accounting isn't 100% accurate' and suggests lowering spark.comet.exec.memoryPool.fraction if you hit out-of-memory errors. Native operators can spill, bounded by spark.comet.maxTempDirectorySize (100 GB per task by default). Comet shuffle is enabled with the CometShuffleManager, in a fully native mode or a JVM-based columnar mode that supports complex key types.

Sources: datafusion.apache.org/comet tuning guide · Fetched 27 Sep 2026

What does the project claim?

Project states

Acceleration over stock Apache Spark on TPC-DS at 1 TB, with a per-query breakdown in the Benchmarking Guide.

Source: datafusion.apache.org/comet · Fetched 27 Sep 2026

Project figure. Not measured by ETL Compare.

Where is it strong?

  • Free, prebuilt and quick to try on Linux (cost model 5.0, tied highest).
  • Current Spark support, including Spark 4.x.
  • Native shuffle and native Parquet scans, including for Iceberg tables.

What are the watch-outs?

  • Memory settings need care; the project says so itself.
  • The write path is not described on the pages we reviewed.
  • Community support only; younger than Gluten and the RAPIDS Accelerator.

Who should shortlist it?

Teams on Spark 3.5 or 4.x that self-manage Spark and want an open-source plugin with a small footprint and a published benchmark.

Which job stages does it address?

Spark ETL job stages and which stages each accelerator documents addressingA waterfall of five Spark ETL job stages (read, transform, shuffle, spill, write) with illustrative proportions, and below it a grid showing, for each accelerator, whether its own documentation says it addresses that stage.Anatomy of a Spark ETL jobRead: scan and decode filesTransform: filter, join, aggregateShuffle: write and fetch between stagesSpill: memory pressure pushes data to diskWrite: encode and commit outputIllustrative proportions, not measured data. Your own split comes from the Spark UI: see Profile a slow Spark job.
Stage coverage by accelerator, from vendor documentation
AcceleratorReadTransformShuffleSpillWrite
CometDocumentedDocumentedDocumentedDocumentedDocumentedDocumentedPartialPartialNot statedNot stated
  • DocumentedDocumented: the vendor's own documentation says it addresses this stage
  • PartialPartial: indirect or experimental, or covered only by an end-to-end claim
  • Not statedNot stated: not found on the pages we reviewed

This shows what each vendor says, not what we measured. Sources are listed on each review.

Native Parquet scans including for Iceberg; native and columnar shuffle; native operators spill within a per-task limit.

Source: Comet tuning guide · Fetched 27 Sep 2026

Frequently asked questions

Is Comet the same as DataFusion?

No. DataFusion is the Rust query engine. Comet is the Spark plugin that runs Spark operators on it.

Does Comet need GPUs or special hardware?

No. The project says it runs on commodity hardware.

Which Spark versions does Comet support?

Spark 3.4.3 (deprecated), 3.5.9, 4.0.4 and 4.1.3, with experimental 4.2.0, per the installation guide.

Compare with

Before you buy

Find your slowest stage with Profile a slow Spark job, then work out the money with Spark cost per job, explained.

Sources