Glossary
What is columnar execution?
In brief
Columnar execution keeps data in memory column by column, so an operation on one column reads only that column's values, stored next to each other. It pairs naturally with columnar file formats such as Parquet and with vectorized engines, and it is the in-memory layout used by Apache Arrow, which DataFusion (under Comet, Auron and Flarion) uses as its in-memory format.
How is it different from columnar storage?
Columnar storage (Parquet, ORC) is how data sits on disk. Columnar execution is how the engine holds and processes it in memory. Spark reads Parquet columnar on disk but, outside accelerated paths, processes much of it as rows in the JVM; native engines keep it columnar through more of the plan.
Why do accelerators care about it?
Converting between rows and columns costs time. Accelerators that keep data columnar from scan through shuffle avoid repeated conversions, which is one reason Gluten and Comet ship their own columnar or native shuffle managers.
Related terms
Sources
- Apache DataFusion · Fetched 27 Sep 2026