Searcharxiv⌕ Search

arXiv · 2610.09793

Formal Runtime Verification for Tool-Using LLM Agents: An Offline Same-Benchmark Study on AgentDojo and STAC

Abstract

Guardrails for tool-using LLM agents are usually application-specific rules, which makes multi-step, data-dependent safety policies hard to specify, audit and reuse. As a declarative alternative, we evaluate metric first-order temporal logic (MFOTL), replaying the recorded trajectories that AgentDojo, STAC and R-Judge already ship through the unmodified MonPoly monitor, offline and without running an agent. On these corpora, five generic obligations flag 71.8% of STAC attack chains and 70.1% of successful AgentDojo attacks, but also fire on 29.3% of benign runs. This imprecision stems from the corpora rather than the logic: they rarely record approvals and never record timestamps, so history-dependent obligations reduce to detecting risky action types. Where the trace does carry relational context, provenance-aware policies discriminate better; that context, however, is itself attackable, and one planted line defeats a naive provenance check on 94-99% of the runs it would otherwise flag. Binding provenance to the lookup that produced it closes this evasion at no cost in detection or benign firing. Taken together, these results show that formal temporal monitoring adds value exactly when the trace exposes trustworthy history. We therefore quantify how far current benchmarks are from that point and propose a twelve-field enforcement-ready trace schema.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Nikolaos Kekatos, Stylianos Basagiannis, Marinelio Chintri, Alexios Lekidis, Tom Nianios, Ioannis Seitoglou, Anastasios Temperekidis, Panagiotis Katsaros. 2026-10-07. Formal Runtime Verification for Tool-Using LLM Agents: An Offline Same-Benchmark Study on AgentDojo and STAC. https://doi.org/10.1145/3847353.3847499

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model

LLM inference is inherently expensive, even a modest slowdown can translate into substantial operating costs and severe availability risks. Recently, a growing body of research known as latency attacks focuses on crafting inputs to trigger worst-case output lengths. However, we report a contrary finding that these algorithmic-level latency attacks are largely ineffective against modern LLM serving systems. We reveal that system-level optimization such as continuous batching provides a logical isolation to mitigate contagious latency impact on co-located users. Thus, in this paper, we shift our focus from the algorithm to the system layer, and introduce a new Fill and Squeeze attack strategy targeting the state transition of the scheduler. ``Fill'' first exhausts the global KV cache to induce Head-of-Line blocking, while ``Squeeze'' forces the system into repetitive preemption. By manipulating output lengths using different attack prompts, and leveraging side-channel probing of memory status, we demonstrate that the attack can succeed in a practical black-box setting with much less cost. Extensive evaluations on vLLM indicate up to $75-742\times$ TTFT degradation relative to benign baselines and $1.5-4\times$ average slowdown on Time Per Output Token compared to existing attacks with 30-40% lower attack cost. Code: https://github.com/Phil-Fan/FS-attack

cs.CR↗

mAVE: A Watermark for Joint Audio-Visual Generation Models

Watermarking joint audio-visual generation supports vendor copyright protection and content provenance. However, independently valid audio and video watermarks do not establish a shared generation session. An adversary can splice watermarked modalities from different sessions, causing the pair to be mistaken for the vendor's original joint output. We introduce mAVE (Manifold Audio-Visual Entanglement), a training-free watermarking framework that strengthens vendor attribution through session binding in native joint audio-visual diffusion transformers. mAVE separates public record retrieval from secret session authentication: a fixed public index locates the server record, while a randomized payload binds audio bits to a session-keyed video grid through a cryptographic digest. One prompt-conditioned joint inversion supports provider-assisted verification of both modalities against a session record, without modifying generator weights or training auxiliary watermark networks. Our analysis establishes implementation-matched distribution preservation and a full-initialization routing/clipping budget, alongside adaptive session-pool security and stable local-perturbation bounds. Experiments on LTX-2 and MOVA show comparable generation quality. mAVE achieves 99.8\% true-positive rate and 0\% observed false-positive rate in the evaluated swap test, and retains 99.2\% true-positive rate under FrameAvg temporal averaging. Same-prompt and similarity-selected swaps further test session authentication beyond perceptual compatibility.

cs.CR↗

Context-Binding Gaps in Stateful Zero-Knowledge Proximity Proofs: Taxonomy, Separation, and Mitigation

A zero-knowledge proximity proof certifies geometric nearness but carries no commitment to an application context. In stateful geo-content systems, where drops can share coordinates, policies evolve, and content has persistent identity, this gap can permit proof transfer between application objects. We present a systems-security analysis of this deployment problem: a taxonomy of context-binding vulnerabilities; a formal model whose replay game asks whether a recorded proof transcript can be re-bound to a different application context (fresh in-radius proving is provably beyond any statement-level mechanism and is delegated to an orthogonal presence layer); an assumption comparison across five binding strategy classes; and a concrete instantiation, Zairn-ZKP, that embeds drop identity, policy version, and session context as public circuit inputs. In-proof binding removes the nonce-to-drop mapping and nonce-uniqueness invariants from the operational assumption set and adds no measurable proving cost over a sound geo-only baseline. A hardened stored-digest check blocks the same transfer attacks under per-request nonces, but its resistance comes from an in-statement challenge digest -- a hybrid, not a purely off-circuit design -- while purely off-circuit strategies cannot resist an adversary able to request fresh challenges; holding nonce policy constant, in-proof context binding is the only strategy blocking same-epoch transfer under shared nonces. Measurements across six network conditions, seven venues in four countries, and an epoch-window simulation indicate same-epoch transfer is a realistic concern in dense urban deployments. Evaluation spans five platforms, seven strategies, and an end-to-end transfer attack; all artifacts are public.

cs.CR↗