ETL CompareETL Compare

Notes · Published 22 September 2026

Spark accelerator news, June to September 2026: what changed

ETL Compare staff · Published 22 September 2026 · Figures checked September 2026

  • News
  • Releases

Four months of releases, published results and technical writing, grouped by theme, with the source for every item.

In brief

This post covers announcements from 3 June to 10 September 2026. The biggest changes were releases: Apache Spark 4.2.0 on 14 July, Comet 1.0.0 on 7 August, NVIDIA cuDF for Apache Spark 26.08.0 on 20 August and Apache Gluten 1.7.0 on 26 August. Alongside them, AWS engineers, the Velox team and DualBird published results, and DualBird and Flarion published technical posts on skew and on Spark upgrades.

What does this recap cover?

It summarises the dated items on our news page between 3 June and 10 September 2026, grouped by theme rather than by date. Every figure below is stated by the vendor, project or authors who published it; we did not measure any of them. For earlier items, including DualBird's funding announcement in February and its April post on why Spark jobs spill, see the news page.

Which releases shipped?

Apache Spark 4.2.0 was released on 14 July 2026. Its release notes highlight GEOMETRY and GEOGRAPHY types, a SQL CHANGES clause for change data capture, Arrow-optimized Python UDFs switched on by default and shuffle checksum validation.

Apache DataFusion Comet had two releases in the period. Comet 0.17.0 (20 June) added a JVM code generation dispatcher and support for Java and Scala UDFs so that fewer queries fall back to Spark, and the project states about 9% faster TPC-DS results from removing round trips between the JVM and native code. Comet 1.0.0 (7 August) supports Spark 3.5.9, 4.0.4 and 4.1.3, with Spark 4.2 experimental, deprecates Spark 3.4 and JDK 11 with removal planned in 1.1.0, and adds Iceberg table format V3 support and experimental PyArrow UDFs. It includes 244 commits from 23 contributors.

NVIDIA cuDF for Apache Spark 26.08.0, from the project formerly called the RAPIDS Accelerator for Apache Spark, is dated 20 August. Its release notes name Spark 3.5.9, 4.0.3, 4.0.4, 4.1.2, 4.1.3 and 4.2.0, add Iceberg 1.11 support with accelerated incremental append scans, and drop Databricks 13.3 ML LTS.

Apache DataFusion 55.0.0 (25 August), the engine under Comet, Auron and Flarion, adds range partitioning, MERGE INTO planner support, runtime row-group pruning and pluggable spill backends, from 877 commits by 175 contributors. Apache Gluten 1.7.0 (26 August) adds native Delta Lake 4 writes for Spark 4.0, native Parquet writes for complex types and a columnar table cache on by default, and updates the project's build and release files for its graduation to an Apache top-level project.

What results were published?

On 3 June, two AWS engineers, Manabu McCloskey and Vara Bonthu, reported on the Apache DataFusion blog that they ran a 3 TB TPC-DS workload on Spark 3.5.8 on Amazon EKS. They state that their first Comet run was 11% slower than vanilla Spark and that, after two months of fixes with the maintainers, Comet was 32% faster overall and 37% faster on Iceberg tables.

On 11 June, DualBird published a sort-based Iceberg compaction benchmark on a 100 GB dataset of about 370 million rows and 64 mixed-type columns, compressed with ZSTD level 3.

VENDOR STATES

"12-20x faster Spark task execution" and "55-85% lower EC2 cost", against vanilla Spark and "state-of-the-art C++ accelerated Spark"

Source: DualBird Iceberg compaction benchmark · Fetched 27 Sep 2026

Vendor figure. Not measured by ETL Compare.

On 25 August, Hongze Zhang (IBM) and Xiaoxuan Meng (Meta) described on the Velox blog how Gluten now collects Delta Lake file statistics natively in Velox. The authors state that Gluten's Delta write went from 2.22% slower than vanilla Spark to 61.31% faster in their test.

The three results use different workloads, data sizes and baselines, so they cannot be compared with each other. Our lesson on reading Spark accelerator benchmarks explains what to check.

What did the commercial vendors publish?

On 1 September, DualBird co-founder and chief architect Ehud Eliaz published the second part of a series explaining how data skew causes spills and straggler tasks in Spark. He states that DualBird's pipeline keeps only a handful of partitions live at once, so each one has a larger memory budget. The post gives no figures.

Flarion's Ran Reichman published two posts. On 18 August he argued that proving output data is unchanged, rather than rewriting code, is the main obstacle in Spark version upgrades, citing published migrations at Slack, Uber and Facebook. On 10 September he argued that Spark will persist because of its scale, network effects and the cost of migrating away, and that its execution layer will be replaced underneath a stable API.

What does it mean for buyers?

  • Spark 4.2 support is uneven. Check the exact version each product lists before an upgrade: see Spark 4.x support in Spark accelerators.
  • Open-source results move quickly. The AWS engineers' Comet numbers went from slower to faster over two months of fixes, which is a reason to plan a tuning round into any trial.
  • Table formats are where much of the activity is: Comet's Iceberg V3 support, cuDF for Apache Spark's Iceberg 1.11 scans, Gluten's native Delta writes and statistics, and DualBird's compaction benchmark all target Iceberg or Delta Lake workloads.
  • Output validation deserves a line in every trial plan, whichever product you test.

None of these announcements changes the scores in this edition, which were checked on 27 September 2026 against the product pages listed on each review.

Sources

Related