arXiv · 2608.14590
Toward Safe LLM Agents: A Survey of Specification, Verification, and Enforcement
Abstract
LLM agents increasingly perform irreversible real-world actions, including database updates, API calls, file operations, and autonomous use of tools. However, no existing system provides formally grounded, task-level safety guarantees for the plans these agents generate. Research remains fragmented across specification, verification, and enforcement, limiting understanding of the strengths and limitations of existing approaches. To address this gap, we conducted a PRISMA 2020 systematic review of 38 studies published between 2022 and 2026 and retrieved from six academic databases. Our analysis reveals four key findings. First, the specification bottleneck remains the primary challenge: natural-language-to-formal translation achieves only 24% to 35% semantic correctness, undermining downstream verification. Second, runtime monitoring is the most mature enforcement strategy, reducing unsafe actions by 40% to 65% in controlled settings, but it does not provide complete safety guarantees. Third, the verifier tax shows that blocking 94% of unsafe actions can still result in less than 5% safe task completion because agents exploit alternative unsafe paths. Finally, no existing approach simultaneously achieves soundness, scalability, semantic correctness, and task-level safety preservation. We contribute a three-level taxonomy, a comparative analysis of existing techniques, a synthesis of evidence on the verifier tax, and a ten-problem research agenda for trustworthy agentic AI.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Pierre Dantas, Lucas Cordeiro, Ehsan Nowroozi, Tihanyi Norbert. 2026-06-22. Toward Safe LLM Agents: A Survey of Specification, Verification, and Enforcement. https://arxiv.org/abs/2608.14590
Cite the original work for its findings. Save a collection to share your selection of sources.