Review · Figures checked September 2026
RAPIDS Accelerator for Apache Spark review (now NVIDIA cuDF for Apache Spark)
ETL Compare staff · Figures checked September 2026 · Sourced from vendor docs, project pages and public benchmarks
In brief
The RAPIDS Accelerator runs supported Spark SQL and DataFrame operations on NVIDIA GPUs with no code changes, and it has the broadest platform list in this set. It scores 3.7 out of 5: strong on stage coverage and maturity, held back by the requirement that every node be a GPU instance and by the GPU-specific tuning that brings.
3.7 / 5
ETL fit score (editorial assessment, 0-5)ETL fit score (editorial assessment, 0-5)
Rank 3 of 7
Best for teams that already run GPU capacity
| Approach | Open-source NVIDIA plugin that runs supported Spark SQL and DataFrame operations on GPUs |
|---|---|
| Runs on | Amazon EMR, Databricks, Dataproc, GKE, Azure Synapse, Kubernetes, on-premises, OCI |
| Instances | NVIDIA GPU instances (Volta or later) |
| License | Apache License 2.0 |
| Price | No plugin fee; GPU instance pricing applies |
| Code changes | None; plugin replaces internal physical plan parts |
| Figures checked | September 2026 |
Sources: spark-rapids overview, RAPIDS Accelerator user guide, RAPIDS Accelerator FAQ, NVIDIA/spark-rapids on GitHub · Fetched 27 Sep 2026
How does it score on each criterion?
| Criterion | Weight | Score | Why |
|---|---|---|---|
| Adoption effort | 20% | 3.4 | No code changes, but the job has to move to NVIDIA GPU instances and the cluster needs GPU-specific configuration. |
| Stage coverage | 20% | 4.2 | Documents GPU execution for group by, joins, sorts and windows, Parquet and ORC writing, CSV reading and a RAPIDS Shuffle Manager; unsupported operations fall back to CPU. |
| Tuning and operating burden | 15% | 3.0 | Qualification and Profiling tools help, but GPU sizing and plugin configuration are your team's job. |
| Platform and instance portability | 10% | 3.2 | The widest platform list in this set (EMR, Databricks, Dataproc, GKE, Azure Synapse, Kubernetes, on-premises, OCI), held back because every one of them needs GPU instances. |
| Cost model transparency | 10% | 4.2 | The plugin is Apache 2.0 with no fee; the cost question becomes GPU instance price against runtime saved, which you can model from public cloud prices. |
| Published evidence | 15% | 3.6 | The Qualification Tool estimates fit from your own event logs; headline benchmark figures were not reviewed for this edition. |
| Maturity and community | 10% | 4.6Tied highest | A long-running NVIDIA project with more than 9,000 commits on main, now published as NVIDIA cuDF for Apache Spark. |
What is the RAPIDS Accelerator?
The RAPIDS Accelerator for Apache Spark is an open-source NVIDIA plugin that uses GPUs to accelerate Spark processing through the RAPIDS libraries, combining cuDF with Spark's distributed execution. Its GitHub repository is now named NVIDIA cuDF for Apache Spark and notes that it was formerly known as the RAPIDS Accelerator; the artifact names are unchanged. The plugin is Apache 2.0 licensed and has more than 9,000 commits on its main branch.
It needs no changes to application code. According to the FAQ, it 'replaces parts of the physical plan that Apache Spark considers internal', works with PySpark, Scala, Java and R, and falls back to the CPU for operations it does not support. Setting spark.rapids.sql.explain to ALL shows which parts of a query will run on the GPU.
Sources: nvidia.github.io/spark-rapids, docs.nvidia.com FAQ, github.com/NVIDIA/spark-rapids · Fetched 27 Sep 2026
Where does it run?
The user guide covers Amazon EMR, Databricks, Google Cloud Dataproc (standard and serverless), Google Kubernetes Engine, Azure Synapse Analytics, Kubernetes, on-premises clusters and Oracle Cloud Infrastructure. On each of them the executors need NVIDIA GPUs; the FAQ refers to Volta or later architectures.
Sources: RAPIDS Accelerator user guide · Fetched 27 Sep 2026
What does it cost?
The plugin itself has no fee. The cost question is whether the runtime saved on GPU instances outweighs their higher hourly price. You can model that from public cloud prices and the Qualification Tool's estimate for your own jobs; see Spark cost per job, explained.
Where is it strong?
- Broadest documented stage coverage among the open-source options: joins, aggregations, sorts, windows, Parquet and ORC writing, CSV reading, and its own shuffle manager.
- The Qualification and Profiling tools analyze your own Spark event logs, which is better evidence for your case than any vendor benchmark.
- Long track record and NVIDIA backing (maturity 4.6, tied highest in set).
What are the watch-outs?
- GPU instances for every executor. That changes instance availability, quotas and price per hour.
- GPU memory, concurrency and shuffle-mode settings are your team's job; the shuffle manager has a default multithreaded mode and a UCX mode with different compatibility.
- The docs list a limited set of accelerated operations; jobs heavy in UDFs or unsupported expressions will fall back to the CPU.
Who should shortlist it?
Teams that already run or can budget GPU capacity, on any major Spark platform, with SQL-heavy ETL. Run the Qualification Tool on a week of event logs before anything else.
Which job stages does it address?
| Accelerator | Read | Transform | Shuffle | Spill | Write |
|---|---|---|---|---|---|
| RAPIDS Accelerator | Documented | Documented | Documented | Not stated | Documented |
- Documented: the vendor's own documentation says it addresses this stage
- Partial: indirect or experimental, or covered only by an end-to-end claim
- Not stated: not found on the pages we reviewed
This shows what each vendor says, not what we measured. Sources are listed on each review.
FAQ lists group by, joins, sorts, window operations, Parquet and ORC writing, CSV reading, and the RAPIDS Shuffle Manager.
Source: RAPIDS Accelerator FAQ · Fetched 27 Sep 2026
Frequently asked questions
Is the RAPIDS Accelerator the same as NVIDIA cuDF for Apache Spark?
Yes. The GitHub repository was renamed to NVIDIA cuDF for Apache Spark and says it was formerly known as the RAPIDS Accelerator for Apache Spark.
Does it need code changes?
No. It replaces internal parts of Spark's physical plan and falls back to the CPU for unsupported operations.
Does it run on Amazon EMR?
Yes, EMR is one of the documented platforms, along with Databricks, Dataproc, Kubernetes and on-premises clusters, all with GPU instances.
Is it free?
The plugin is Apache 2.0 licensed with no fee. You pay for the GPU instances.
Compare with
- DualBird vs RAPIDS Accelerator
- RAPIDS Accelerator alternatives
- Apache Auron vs RAPIDS Accelerator
- DataFusion Comet vs RAPIDS Accelerator
- Flarion vs RAPIDS Accelerator
- Apache Gluten vs RAPIDS Accelerator
- Databricks Photon vs RAPIDS Accelerator
Before you buy
Find your slowest stage with Profile a slow Spark job, then work out the money with Spark cost per job, explained.
Sources
- RAPIDS Accelerator for Apache Spark · Fetched 27 Sep 2026
- RAPIDS Accelerator for Apache Spark FAQ · Fetched 27 Sep 2026
- NVIDIA cuDF for Apache Spark on GitHub · Fetched 27 Sep 2026
- RAPIDS Accelerator for Apache Spark user guide · Fetched 27 Sep 2026