ETL CompareETL Compare

Guide · Figures checked September 2026

Where Spark accelerators run: EMR, Databricks, Dataproc, Glue and open-source Spark

ETL Compare staff · Figures checked September 2026

In brief

Platforms and accelerators are different layers. Amazon EMR, Databricks, Google Cloud's managed Spark service, AWS Glue and self-managed Spark run your jobs; accelerators change how the jobs execute on them. You do not need to leave your platform to use most accelerators: DualBird, the RAPIDS Accelerator and Flarion document Amazon EMR, and the open-source plugins can be added wherever you control Spark configuration.

Platform layer and accelerator layerYour ETL code: Spark SQL, DataFrames, PySpark (unchanged)Accelerator layer: Gluten, Comet, RAPIDS Accelerator, DualBird, Auron, Photon, FlarionPlatform layer: Amazon EMR · Databricks · Google Cloud managed Spark · AWS Glue · self-managed Spark and Kubernetes
Figure 1. Accelerators replace parts of physical-plan execution. The platform stays.

What does each platform provide?

Amazon EMR

AWS's managed big data service. It runs on EC2, on EKS or serverless, and ships the EMR runtime for Apache Spark, which AWS states is up to 5.4x faster than open-source Apache Spark while keeping API compatibility. EMR's Spark optimizations include adaptive query execution, dynamic partition pruning and Bloom filter joins.

Source: Amazon EMR · Fetched 27 Sep 2026

Databricks

A managed data platform on AWS, Azure and Google Cloud. Its runtime includes Photon, the one accelerator in this guide that only exists inside a platform.

Source: Databricks · Fetched 27 Sep 2026

Google Cloud Managed Service for Apache Spark (formerly Dataproc)

Google Cloud's managed Spark clusters; Google now calls the cluster offering Managed Service for Apache Spark.

Source: Google Cloud Managed Service for Apache Spark (formerly Dataproc) · Fetched 27 Sep 2026

AWS Glue

A serverless data integration service whose ETL engine is based on Apache Spark. Because it is serverless, you control less of the cluster, and none of the accelerator pages we reviewed list Glue.

Source: AWS Glue · Fetched 27 Sep 2026

Self-managed Apache Spark (VMs or Kubernetes, including Amazon EKS)

Open-source Spark you deploy yourself. This is where open-source plugins such as Gluten, Comet and Auron are documented to run.

Source: Self-managed Apache Spark (VMs or Kubernetes, including Amazon EKS) · Fetched 27 Sep 2026

These platforms are not scored or ranked on ETL Compare. Choosing a platform involves far more than ETL speed, and this site does not cover it.

Which accelerators are documented for which platform?

Platform support by accelerator, from vendor and project pages
AcceleratorEMRDatabricksGoogle Cloud (Dataproc)AWS GlueSelf-managed Spark / Kubernetes
Gluten + VeloxNot documented(self-install)Not documentedNot documented(self-install)Not documentedDocumented
CometNot documented(self-install)Not documentedNot documented(self-install)Not documentedDocumented
RAPIDS AcceleratorDocumentedDocumentedDocumentedNot documentedDocumented
DualBirdDocumentedNot documentedNot documented(AWS only)Not documentedDocumented(Apache Spark, Amazon EKS)
AuronNot documented(self-install)Not documentedNot documented(self-install)Not documentedDocumented
PhotonNot availableOnly platformNot availableNot availableNot available
FlarionDocumentedDocumentedDocumentedNot documentedDocumented(on-premises)

Documented means the vendor or project lists the platform on the pages we reviewed. Self-install means you can usually add an open-source plugin to a platform that lets you set Spark configuration and classpath, but the project does not publish a guide for that platform. None of the vendor pages we reviewed list AWS Glue.

Do I have to leave Amazon EMR to use an accelerator?

No. On the pages we reviewed, three accelerators document Amazon EMR directly: DualBird (which also documents Amazon EKS), the RAPIDS Accelerator and Flarion. Gluten, Comet and Auron do not publish EMR guides, but they are Spark plugins configured through Spark settings and classpath, which EMR lets you set. The only accelerator here that requires a different platform is Databricks Photon.

Why is EMR's own runtime not in the ranking?

Because it is part of the platform. The EMR runtime is the baseline an EMR team starts from. When a vendor quotes a speedup, ask whether it was measured against open-source Spark or against the EMR runtime; the answer changes what the figure means for you.

Frequently asked questions

Is Amazon EMR a competitor to Spark accelerators?

No. EMR is a platform that runs Spark. Accelerators run on top of it, and several document EMR support.

Which accelerators work on Databricks?

Photon is built in. The RAPIDS Accelerator and Flarion also list Databricks.

Can I use an accelerator with AWS Glue?

None of the accelerator pages we reviewed list AWS Glue.

Related