Searcharxiv⌕ Search

arXiv · 2609.32414

How AI Changes DevOps Performance: A Mechanism-Based Simulation

Abstract

Context. AI tools for software development have progressed from code completion to autonomous agents, yet evidence on their effects on DORA performance is mixed, mostly cross-sectional, and often conflates AI-generated code (AGC) with agentic AI (AGT). Objective. We examine how AGC and AGT, separately and jointly, affect deployment frequency, lead time, change failure rate, recovery time, and deployment rework rate. Method. We audit the evidence, formalise a two-construct model with four capability moderators, and implement a discrete-event simulation of a delivery pipeline. Six experiments include a factorial design, Shapley decompositions, sensitivity analysis over 600 parameter draws, a fixed-demand robustness test, and a staggered-adoption panel of 5,000 simulated teams with counterfactuals. Results. AGC increased failure rate, rework, and recovery time in both capability profiles. When AI increased change volume, AGC shortened lead time only where review capacity was spare (-13%) and lengthened it once review saturated; in the low-capability team, its deployment-frequency gain came entirely from unplanned rework. AGT improved deployment frequency (+31% to +34%), lead time (-8% to -22%), and recovery time (-29% to -38%), but its stability effect depended on autonomous CI repair masking defects, and escaped defects increased in the high-capability team. Capability reduced AGC's absolute failure-rate penalty (7.9 vs. 1.5 points) but not its relative penalty (+34% vs. +84%). Two-way fixed effects underestimated adoption effects by 8-16%; detecting the modelled stability effect required about 50 teams, and its capability moderation about 400. Conclusions. Within the model, DORA metrics respond to AI through queueing, batching, and oversight, not only code quality. We derive measurement rules, guardrails for agentic remediation, and a sample-size-informed field-validation protocol.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mamdouh Alenezi. 2026-09-26. How AI Changes DevOps Performance: A Mechanism-Based Simulation. https://arxiv.org/abs/2609.32414

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Is Agent Code Less Maintainable Than Human Code?

Maintainability is a core dimension of software engineering, shaping how code is written, reviewed, and developed over time. While coding agents have demonstrated strong performance on single-issue tasks, it remains unclear how maintainable their code is when future agents build on top of it, potentially leading to compounding downstream effects. We investigate how agent code compares to human code in these maintenance settings, presenting CodeThread, a framework to construct controlled experiments from repository-level coding benchmarks. Applying CodeThread to four frontier coding agents and four benchmarks, we find that agents are less effective at resolving tasks when building on agent code compared to human code, with task resolve rate drops of up to 13.1%. Regression analysis reveals that many traditional software engineering maintainability metrics do not explain this difference. Instead, the clearest signals are subtler behavioral differences in agent code, such as changes to input validation and error handling, along with differences in downstream code size and task difficulty. These findings highlight the need to evaluate these systems not only by immediate task resolution but also by code maintainability, and point to potential sources of downstream errors introduced by agent code.

cs.SE↗

CURATE: Leveraging LLM Agents to Compose, Catalog, and Deploy Reproducible Workflows

Agentic code generation has the potential to accelerate the development of computational workflows while also reducing barriers to entry. However, a key gap remains: existing coding agents focus on code generation and do not address the entire workflow lifecycle, including deployment and sharing. As a result, users develop and stitch modules independently while managing deployment on their own. To address this gap, we propose CURATE (Composition, User-in-the-loop, Reuse, and Automated Task Execution), a novel human-in-the-loop multi-agent system that uses LLM agents to manage and develop composable workflows across their entire lifecycle. A key feature of the system is a catalog that allows for the storage and reuse of modules across workflows. Module catalogs provide a foundation that can be expanded to support FAIR principles by facilitating the sharing and reuse of curated modules and subgraphs. We demonstrate the feasibility of our system with an initial prototype and 6 experiments.

cs.SE↗

MCP Error Messages Written for Developers Hurt the Most Capable Agents Most

Many Model Context Protocol (MCP) servers wrap web APIs built for human developers, and their error messages tell the reader to run a command, edit a configuration, open a web page or wait. Many agents that read them can only call the server's tools. In 150 widely used MCP servers, 949 of 3,001 error messages tell the caller what to do next, and half of these steps depend on something the server cannot see about the caller. On credential errors, 62 of 67 steps ask for a terminal command, a configuration change or a web page; on rate limits, 20 of 30 say to wait and retry without naming the call to repeat. We tested five OpenAI models that act only through the tools of Berkeley Function Calling Leaderboard tasks, and the agents did what the step said. On expired credentials, a terminal command in the step left 45% of tasks recovered, and the loss it caused grew from 18 points for GPT-5.5 to 69 for GPT-6 Astra. On a rate limit, GitHub's "Wait before retrying." left 6%. We tested two remedies. For MCP developers, naming a server tool in the step raised recovery on expired credentials to 84%, with the login tool in place of the command, and on a rate limit to 88%, with the call to repeat in place of the bare wait. For agent developers, deleting the step with a one-sentence prompt before the model reads it raised recovery on expired credentials to 82%.

cs.SE↗