arXiv · 2609.37475
Boundary-State Control for Tool-Using Language-Model Agents: Commit-Time Consistency under State Drift
Abstract
Tool-using language-model agents can decide that an action is permissible and execute it only after security-relevant state has changed. We study this proposal-to-commit gap and introduce BSC-R, a deterministic effect-boundary mechanism that binds a single-use commit authorization to the exact action and to a semantic projection of the authorization state that justified it. On 2,847 attacked AgentDojo episodes, the boundary kernel preserves the unprotected agent's behavior exactly (80.576% utility; 1.616% attack success). On 10,302 frozen proposals, it accepts every unchanged commit and rejects every instance of ten prospectively specified changed or replayed classes. In an independently generated boundary-drift experiment, full joint binding commits 0/4,403 invalid contexts while retaining 5,899/5,899 valid contexts. A prospective external evaluation on the 3,460-scenario CONTINUITY suite retains all 700 benign cases, prevents 1,200/1,200 represented non-replay invalid effects, handles 160/160 replay lifecycles correctly, and withholds 200/200 ambiguous no-release cases. The broader external suite also exposes the method's limit: across all 2,560 attacks, BSC-R has a 25% invalid-effect commit rate versus 0% for CONTINUITY. The result is therefore a scoped commit-time consistency mechanism, not a universal agent-safety claim.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wesley Shu. 2026-09-26. Boundary-State Control for Tool-Using Language-Model Agents: Commit-Time Consistency under State Drift. https://arxiv.org/abs/2609.37475
Cite the original work for its findings. Save a collection to share your selection of sources.