ETL CompareETL Compare

Review · Figures checked September 2026

Databricks Photon review: native execution inside Databricks

ETL Compare staff · Figures checked September 2026 · Sourced from vendor docs, project pages and public benchmarks

In brief

Photon is the easiest accelerator to adopt if your jobs already run on Databricks: it is on by default for serverless and SQL warehouses and a single setting elsewhere, with no plugin to maintain. It scores 3.4 out of 5 on this site because the site's question is speeding up Spark ETL where it runs today, and Photon only exists inside Databricks.

3.4 / 5

ETL fit score (editorial assessment, 0-5)

ETL fit score (editorial assessment, 0-5)

Rank 6 of 7

Best for teams already on Databricks

Photon at a glance
ApproachDatabricks-native vectorized C++ query engine inside the Databricks Runtime
Runs onDatabricks only (AWS, Azure, Google Cloud)
InstancesSupported Databricks instance types, including Graviton
LicenseCommercial (part of Databricks)
PriceDatabricks DBU pricing; Photon changes DBU consumption
Code changesNone on Databricks; platform move otherwise
Figures checkedSeptember 2026

Sources: Photon documentation · Fetched 27 Sep 2026

How does it score on each criterion?

Score breakdown by criterion, with reasons
CriterionWeightScoreWhy
Adoption effort20%2.6Lowest in setOne checkbox (or on by default on serverless) with no code changes if you already run on Databricks; for a team on EMR or self-managed Spark it means moving to Databricks first.
Stage coverage20%3.8Documents native scans (Parquet, Delta, CSV, JSON), filters, joins, aggregates, windows and writes, with fallback to the Spark runtime; no UDF, RDD or stateful streaming support.
Tuning and operating burden15%4.6Highest in setManaged by Databricks inside its runtime, so there is no plugin for your team to install or version.
Platform and instance portability10%1.6Lowest in setAvailable only inside Databricks (on AWS, Azure and Google Cloud).
Cost model transparency10%3.2Databricks publishes DBU rates, but Photon compute consumes DBUs at different rates, so the net effect has to be modelled per job.
Published evidence15%3.2Databricks documents where Photon does not help (queries under two seconds, UDFs); performance comparisons outside Databricks are not possible because it only runs there.
Maturity and community10%4.6Tied highestThe default engine on Databricks serverless and SQL warehouses, backed by Databricks.

What is Photon?

Databricks describes Photon as 'the Databricks-native vectorized query engine that accelerates your SQL workloads, DataFrame API calls, ETL pipelines, and stateless streaming workloads'. It replaces JVM-based execution with native C++ that processes data in columnar batches, and it is compatible with Apache Spark APIs, so existing code runs without changes.

Sources: docs.databricks.com/aws/en/compute/photon · Fetched 27 Sep 2026

How do you turn it on?

On serverless compute, SQL warehouses and serverless Lakeflow pipelines, Photon is on by default. On classic all-purpose and jobs compute you tick 'Use Photon Acceleration' or set runtime_engine to PHOTON in the API.

What does it accelerate, and where does it fall back?

The documentation lists scans of Parquet, Delta, CSV and JSON, filters, joins, aggregates, window functions and writes. For unsupported operations Photon 'transparently falls back to the Spark runtime for that portion of the execution'. Databricks also states that Photon does not improve queries that finish in under two seconds and does not support UDFs, stateful streaming or RDD APIs.

What does it cost?

Photon compute consumes DBUs at different rates from non-Photon compute, so the saving depends on whether the runtime reduction outweighs the rate change for each job. Databricks' pricing pages give the rates; the arithmetic is in Spark cost per job, explained.

Where is it strong?

  • Lowest operating burden in the set (4.6): nothing to install, version or patch.
  • Mature and backed by Databricks (maturity 4.6, tied highest).
  • Clear documentation of what it does not accelerate, which saves wasted trials.

What are the watch-outs?

  • Databricks only (portability 1.6, lowest in set). For a team on EMR or self-managed Spark, adopting Photon means a platform migration, which is why adoption effort scores 2.6, the lowest here.
  • UDF-heavy and RDD-based ETL will not benefit.
  • DBU rate changes can offset runtime gains on short jobs.

Who should shortlist it?

Any team already on Databricks should test Photon on its ETL jobs before looking at anything else. Teams elsewhere should not move platforms for Photon alone without costing the migration.

Which job stages does it address?

Spark ETL job stages and which stages each accelerator documents addressingA waterfall of five Spark ETL job stages (read, transform, shuffle, spill, write) with illustrative proportions, and below it a grid showing, for each accelerator, whether its own documentation says it addresses that stage.Anatomy of a Spark ETL jobRead: scan and decode filesTransform: filter, join, aggregateShuffle: write and fetch between stagesSpill: memory pressure pushes data to diskWrite: encode and commit outputIllustrative proportions, not measured data. Your own split comes from the Spark UI: see Profile a slow Spark job.
Stage coverage by accelerator, from vendor documentation
AcceleratorReadTransformShuffleSpillWrite
PhotonDocumentedDocumentedDocumentedDocumentedNot statedNot statedNot statedNot statedDocumentedDocumented
  • DocumentedDocumented: the vendor's own documentation says it addresses this stage
  • PartialPartial: indirect or experimental, or covered only by an end-to-end claim
  • Not statedNot stated: not found on the pages we reviewed

This shows what each vendor says, not what we measured. Sources are listed on each review.

Docs list scans (Parquet, Delta, CSV, JSON), filtering, joins, aggregates, window functions and writes.

Source: Photon documentation · Fetched 27 Sep 2026

Frequently asked questions

Does Photon require code changes?

No. Databricks states Photon is compatible with Apache Spark APIs.

Can I use Photon on Amazon EMR?

No. Photon is part of the Databricks Runtime and is not available outside Databricks.

Does Photon speed up UDFs?

No. Databricks lists UDFs, stateful streaming and RDD APIs as unsupported; those parts run on the Spark runtime.

Compare with

Before you buy

Find your slowest stage with Profile a slow Spark job, then work out the money with Spark cost per job, explained.

Sources