University · Running it · Lesson 8
Spark event logs and the History Server: keeping evidence from every run
ETL Compare staff · Published 14 September 2026 · 3 min read
In brief
Turn on Spark event logging for every run you want to compare. The event log records what the Spark UI shows, the History Server rebuilds that UI after the job ends, and its REST API returns the same data as JSON. With logs for both baseline and trial runs, you can compare stage by stage instead of relying on wall-clock time alone.
What is a Spark event log?
An event log is a file Spark writes as an application runs, recording job, stage and task events with their metrics. The Spark monitoring documentation says that setting spark.eventLog.enabled to true "configures Spark to log Spark events that encode the information displayed in the UI to persisted storage". spark.eventLog.dir sets where the logs go. Point it at storage that outlives the cluster, such as object storage, so the logs survive when a cluster is shut down.
How do you read a finished job?
The History Server rebuilds the web UI of finished applications from their event logs. It listens on port 18080 by default. Its REST API, under /api/v1, returns the same data as JSON, which is the practical way to pull stage metrics for many runs into a spreadsheet or notebook instead of reading them page by page.
What should you keep for each run?
- The event log file itself
- The Spark properties from the Environment tab, including every accelerator setting
- Cluster size, instance type and hourly price
- The input data version, and output row counts or checksums
- Wall-clock runtime and billed cluster time
Keep them together, per run, under one run label. A proof of concept is only as good as its records: without the baseline's event log you cannot tell whether a trial run was faster because of the product or because the input was smaller that day.
Do accelerators use the event log?
Some do. NVIDIA's Qualification Tool for the RAPIDS Accelerator reads event logs from CPU runs and estimates which applications are good candidates for GPU acceleration, and NVIDIA documents a Profiling Tool alongside it. For the other products, the event log is your own evidence: compare the same stages across baseline and trial runs, and check the query plans in the SQL tab to see which operators ran in the accelerator and which fell back to standard Spark.
How long should you keep the logs?
At least until the purchase decision and the first months in production are behind you. Event logs let you rerun a comparison when a vendor ships a new release or when Spark itself moves, as it did with 4.2.0 on 14 July 2026. They are also what a vendor's or project's engineers will ask for when a trial result looks wrong.
Related
Sources
- Monitoring and instrumentation - Spark documentation · Fetched 27 Sep 2026
- Web UI - Spark documentation · Fetched 27 Sep 2026
- RAPIDS Accelerator for Apache Spark user guide · Fetched 27 Sep 2026
- Apache Spark 4.2.0 release · Fetched 27 Sep 2026