ETL CompareETL Compare

News · Updated September 2026

Spark accelerator news

Summaries of public announcements, with links to the original source. Updated monthly.

In brief

The main 2026 developments so far: Apache Spark 4.2.0 shipped on 14 July, Apache DataFusion Comet reached 1.0.0 on 7 August, and Apache Gluten released 1.7.0 on 26 August. The RAPIDS Accelerator added Spark 4.2.0 support in August, and the commercial vendors DualBird and Flarion published technical posts, with DualBird also announcing a funding round and a compaction benchmark.

Flarion

Flarion argues Spark will keep its API while its engine is replaced

Flarion's Ran Reichman writes that Spark remains the main data processing platform because of its scale, network effects and the cost of migrating away, and expects its execution layer to be replaced underneath a stable API.

Source: Flarion blog

DualBird

DualBird explains how its pipeline handles skew-driven spill

In the second part of its "Faster Spark" series, DualBird co-founder and chief architect Ehud Eliaz explains how data skew causes spills and straggler tasks in Spark. He states that DualBird keeps only a handful of partitions live at once, so each one has a larger memory budget.

Source: DualBird blog

Gluten

Apache Gluten 1.7.0 released

The release adds native Delta Lake 4 writes for Spark 4.0, native Parquet writes for complex types and a columnar table cache on by default, and updates the project's build and release files for its graduation to an Apache top-level project.

Source: GitHub: apache/incubator-gluten

Gluten

Gluten moves Delta statistics collection into Velox

Engineers from IBM and Meta describe collecting Delta Lake file statistics natively in Velox instead of row by row in the JVM. The authors state Gluten's Delta write went from 2.22% slower than vanilla Spark to 61.31% faster in their test.

Source: Velox blog

DataFusion

Apache DataFusion 55.0.0 released

The engine under Comet, Auron and Flarion adds range partitioning, MERGE INTO planning, runtime row-group pruning and pluggable spill backends, from 877 commits by 175 contributors.

Source: Apache DataFusion blog

RAPIDS Accelerator

NVIDIA cuDF for Apache Spark 26.08.0 adds Spark 4.2.0

The release notes for the project formerly called the RAPIDS Accelerator for Apache Spark name Spark 3.5.9, 4.0.3, 4.0.4, 4.1.2, 4.1.3 and 4.2.0, and add Iceberg 1.11 support with accelerated incremental append scans.

Source: NVIDIA cuDF for Apache Spark docs

Flarion

Flarion: validating output is the hard part of Spark upgrades

Ran Reichman argues that proving output data is unchanged, rather than rewriting code, is the main obstacle in Spark version upgrades, citing published migrations at Slack, Uber and Facebook.

Source: Flarion blog

Comet

Apache DataFusion Comet reaches 1.0.0

Comet 1.0.0 supports Spark 3.5.9, 4.0.4 and 4.1.3, with Spark 4.2 experimental, deprecates Spark 3.4 and JDK 11, and adds Iceberg table format V3 support and experimental PyArrow UDFs. It includes 244 commits from 23 contributors.

Source: Apache DataFusion blog

Apache Spark

Apache Spark 4.2.0 released

The release highlights GEOMETRY and GEOGRAPHY types, a SQL CHANGES clause for change data capture, Arrow-optimized Python UDFs on by default and shuffle checksum validation.

Source: Apache Spark

Comet

Comet 0.17.0 keeps more queries native

Comet adds a JVM code generation dispatcher and support for Java and Scala UDFs so fewer queries fall back to Spark. The project states about 9% faster TPC-DS results from removing round trips between the JVM and native code.

Source: Apache DataFusion blog

DualBird

DualBird publishes an Iceberg compaction benchmark

The benchmark covers sort-based compaction of a 100 GB dataset of about 370 million rows and 64 columns. DualBird states "12-20x faster Spark task execution" and "55-85% lower EC2 cost" against vanilla Spark and a C++ accelerated Spark baseline.

Source: DualBird blog

Comet

AWS engineers report Comet results on Amazon EKS

Two AWS engineers ran a 3 TB TPC-DS workload on Spark 3.5.8 on Amazon EKS. They report a first Comet run 11% slower than vanilla Spark and, after two months of fixes with the maintainers, 32% faster overall and 37% faster on Iceberg tables.

Source: Apache DataFusion blog

DualBird

DualBird explains why Spark jobs spill

Ehud Eliaz writes that disk spills come from tasks sharing executor memory, and states that DualBird runs a handful of tasks at a time, each with more memory, which it says reduces spill.

Source: DualBird blog

DualBird

DualBird raises $25M led by Lightspeed Venture Partners

DualBird announced $25 million in funding led by Lightspeed Venture Partners, with Angular Ventures, Uncork Capital and Bessemer Venture Partners. It said the money will build out sales and customer success, expand go-to-market and strengthen enterprise partnerships ahead of general availability in early 2026.

Source: DualBird blog

Vendor and project figures in these summaries are stated by the vendor or project and are not measured by ETL Compare.

More analysis in Notes