arXiv · 2609.35199
Trajectory-Level Security Debt in LLM Coding Agents
Abstract
LLM coding agents can traverse hundreds of intermediate code states before submitting a solution. Evaluating only the final artifact leaves the evolution of security findings unmeasured. We introduce the Security Debt Line Integral (SDLI), which accumulates static-analysis risk when an agent reaches a new best test pass ratio. We instantiate it with four static application security testing (SAST) tools and study artifacts from 830 passing SWE-bench runs, 712 ProgramBench final workspaces, and 13 public MirrorCode trajectories. The two large populations use the final-state special case of SDLI. Two-tool Common Weakness Enumeration (CWE) class agreement occurs in 3.9% of SWE-bench runs and 26.2% of the 80 ProgramBench runs passing at least 90% of official tests. These are scanner findings, not validated vulnerability rates. Excluding three advisory-heavy classes reduces the latter rate to 6.2%. Same-task runs differ in their measured scores, while one reconstructed ProgramBench run exposes persistent findings from its first implementation write. A repair case study reduces the scanner signal while preserving tested behavior, but also reveals sensitivity to equivalent API rewrites. SDLI offers a way to study progress and security findings together. Its value for steering agents and confirming exploitable vulnerabilities remains to be established.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Prateek Kumar Rajput, Abdoul Kader Kabore, Yewei Song, Melissa Tessa, Tailia Malloy, Jacques Klein, Tegawendé F. Bissyandé. 2026-09-28. Trajectory-Level Security Debt in LLM Coding Agents. https://arxiv.org/abs/2609.35199
Cite the original work for its findings. Save a collection to share your selection of sources.