arXiv · 2609.33763
SecProbe: Adaptive Evaluation of Coding Agents on Cybersecurity Vulnerabilities
Abstract
Assessing cybersecurity vulnerability awareness in coding agents requires evaluations that reveal capability gaps and remain informative as models evolve. Static benchmarks offer fixed coverage and difficulty, while scarce vulnerable repositories and costly expert authoring limit their renewal at scale. We introduce SecProbe, a framework for adaptive evaluation that combines Item Response Theory (IRT) with on-demand synthesis of repository-scale vulnerability-repair tasks. From observed performance, \textsc{SecProbe} estimates agent ability and identifies where additional evidence is most informative, selecting existing tasks or synthesizing new ones accordingly. As one use case, we construct 353 tasks spanning six programming languages and 151 CWE types and evaluate nine frontier models with two agent harnesses. Success rates peak at 28.33\%, highlighting substantial gaps in vulnerability recognition and repair. Compared with random and one-shot baselines, \textsc{SecProbe} achieves comparable agent ability estimates while requiring agents to solve up to 29.5\% fewer tasks. These results support adaptive evaluation as an efficient and discriminative approach to assessing cybersecurity vulnerability awareness.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xiaonan Luo, Yue Huang, Kehan Guo, Ping He, Chuan Zou, Chujie Gao, Lichi Li, Yuchen Ma, Zhangchen Xu, Zichen Chen, Yufei Han, Xiangliang Zhang. 2026-09-27. SecProbe: Adaptive Evaluation of Coding Agents on Cybersecurity Vulnerabilities. https://arxiv.org/abs/2609.33763
Cite the original work for its findings. Save a collection to share your selection of sources.