arXiv · 2609.33023
SRE-Marathon: A Continuous, Change-Driven Benchmark for Autonomous Site Reliability Agents
Abstract
Benchmarks for site reliability engineering (SRE) agents are typically episodic: one fault is injected, the agent receives an incident task, and its response is scored. Production operation is not. Incidents surface through noisy alerts, overlap in time, and often originate from code or configuration changes. We present SRE-Marathon, a benchmark for long-horizon, continuous SRE operation. An agent is invoked at a fixed cadence with cumulative alert history and a persistent workspace while operating a live two-zone Kubernetes deployment as a fault orchestrator injects overlapping faults according to a seeded, production-calibrated schedule. Curated code and configuration changes deployed through the same build pipeline available for repair. Each run is recorded into a sealed bundle and scored offline: Marathon-Score credits each injected fault for ordered progress through correlation, localization, and repair, with all metrics computed deterministically from recorded system evidence. Across three applications and about sixty faults per run, the best of 10 methods reaches only 41.3 out of 100. Agents often correlate and localize faults, but almost never complete repairs while the faults remain active.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yifang Tian, Yingjian Bai, Yifeng He, Zichun Chong, Yuanchen Gao, Yiran Li, Hans-Arno Jacobsen. 2026-09-26. SRE-Marathon: A Continuous, Change-Driven Benchmark for Autonomous Site Reliability Agents. https://arxiv.org/abs/2609.33023
Cite the original work for its findings. Save a collection to share your selection of sources.