Glossary
What is vectorized execution?
In brief
Vectorized execution means an engine applies each operation to a batch of values at once instead of one row at a time, which cuts per-row overhead and lets the CPU use SIMD instructions. Photon, Velox (used by Gluten) and DataFusion (used by Comet and Auron) are vectorized engines that replace parts of Spark's row-at-a-time JVM execution.
Why does it make Spark faster?
Row-at-a-time execution pays function-call and object overhead for every row. Batch execution pays it once per batch and keeps data in contiguous memory, which suits modern CPUs. Databricks describes Photon as processing data in columnar batches in native C++; Auron describes vectorized computation on Apache Arrow for SIMD-enabled batch processing.
What does it not fix?
Vectorization speeds up computation inside a task. It does not remove a shuffle the query plan needs, and it does not by itself stop a task from spilling if its partition is too large for its memory. Profile first.
Related terms
Sources
- What is Photon? (Databricks documentation) · Fetched 27 Sep 2026
- Apache Auron · Fetched 27 Sep 2026