Searcharxiv⌕ Search

arXiv · 2610.09469

Secure-CUA: Controlling Untrusted Influence in Computer-Use Agents

Abstract

Computer-use agents (CUAs) perform tasks across applications (such as desktops, mobile apps, and web browsers) by observing graphical interfaces and issuing commands such as clicks and keystrokes. These interfaces combine trusted controls and content with untrusted content needed for legitimate tasks. An adversary controlling this untrusted content can embed instructions or misleading visual cues to change the agent's intended action or redirect its commands to the wrong interface target. We formalize security requirements for both the agent's decisions and their execution through GUI commands. In an ideal execution model, we show that enforcing both requirements at each step protects execution traces. We instantiate this model in Secure-CUA, our system for secure CUA execution. Its key idea is to commit to an explicit per-action program, called an $\textit{action transaction}$, before accessing untrusted content. Each transaction fixes its queries to untrusted content and the permitted uses of their responses. The system masks untrusted regions and evaluates each transaction to produce the next action, using an isolated query model to answer its queries. It then locates the intended interface target using the masked interface. Under the model's assumptions, Secure-CUA is secure by design, while generating a new transaction at each step helps maintain high task utility by adapting to changing interfaces. We evaluate Secure-CUA under benign conditions on 400 WebArena tasks using three frontier models across $5$ seeds, yielding $6,000$ execution traces. Secure-CUA achieves an average task success rate of $53.55\%$, compared with $55.12\%$ for Vanilla-CUA and $13.17\%$ for CaMeL-CUA.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sarthak Choudhary, Mihai Christodorescu, Ashish Hooda, Somesh Jha, Tongxin Li, Damien Octeau. 2026-10-07. Secure-CUA: Controlling Untrusted Influence in Computer-Use Agents. https://arxiv.org/abs/2610.09469

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model

LLM inference is inherently expensive, even a modest slowdown can translate into substantial operating costs and severe availability risks. Recently, a growing body of research known as latency attacks focuses on crafting inputs to trigger worst-case output lengths. However, we report a contrary finding that these algorithmic-level latency attacks are largely ineffective against modern LLM serving systems. We reveal that system-level optimization such as continuous batching provides a logical isolation to mitigate contagious latency impact on co-located users. Thus, in this paper, we shift our focus from the algorithm to the system layer, and introduce a new Fill and Squeeze attack strategy targeting the state transition of the scheduler. ``Fill'' first exhausts the global KV cache to induce Head-of-Line blocking, while ``Squeeze'' forces the system into repetitive preemption. By manipulating output lengths using different attack prompts, and leveraging side-channel probing of memory status, we demonstrate that the attack can succeed in a practical black-box setting with much less cost. Extensive evaluations on vLLM indicate up to $75-742\times$ TTFT degradation relative to benign baselines and $1.5-4\times$ average slowdown on Time Per Output Token compared to existing attacks with 30-40% lower attack cost. Code: https://github.com/Phil-Fan/FS-attack

cs.CR↗

mAVE: A Watermark for Joint Audio-Visual Generation Models

Watermarking joint audio-visual generation supports vendor copyright protection and content provenance. However, independently valid audio and video watermarks do not establish a shared generation session. An adversary can splice watermarked modalities from different sessions, causing the pair to be mistaken for the vendor's original joint output. We introduce mAVE (Manifold Audio-Visual Entanglement), a training-free watermarking framework that strengthens vendor attribution through session binding in native joint audio-visual diffusion transformers. mAVE separates public record retrieval from secret session authentication: a fixed public index locates the server record, while a randomized payload binds audio bits to a session-keyed video grid through a cryptographic digest. One prompt-conditioned joint inversion supports provider-assisted verification of both modalities against a session record, without modifying generator weights or training auxiliary watermark networks. Our analysis establishes implementation-matched distribution preservation and a full-initialization routing/clipping budget, alongside adaptive session-pool security and stable local-perturbation bounds. Experiments on LTX-2 and MOVA show comparable generation quality. mAVE achieves 99.8\% true-positive rate and 0\% observed false-positive rate in the evaluated swap test, and retains 99.2\% true-positive rate under FrameAvg temporal averaging. Same-prompt and similarity-selected swaps further test session authentication beyond perceptual compatibility.

cs.CR↗

Context-Binding Gaps in Stateful Zero-Knowledge Proximity Proofs: Taxonomy, Separation, and Mitigation

A zero-knowledge proximity proof certifies geometric nearness but carries no commitment to an application context. In stateful geo-content systems, where drops can share coordinates, policies evolve, and content has persistent identity, this gap can permit proof transfer between application objects. We present a systems-security analysis of this deployment problem: a taxonomy of context-binding vulnerabilities; a formal model whose replay game asks whether a recorded proof transcript can be re-bound to a different application context (fresh in-radius proving is provably beyond any statement-level mechanism and is delegated to an orthogonal presence layer); an assumption comparison across five binding strategy classes; and a concrete instantiation, Zairn-ZKP, that embeds drop identity, policy version, and session context as public circuit inputs. In-proof binding removes the nonce-to-drop mapping and nonce-uniqueness invariants from the operational assumption set and adds no measurable proving cost over a sound geo-only baseline. A hardened stored-digest check blocks the same transfer attacks under per-request nonces, but its resistance comes from an in-statement challenge digest -- a hybrid, not a purely off-circuit design -- while purely off-circuit strategies cannot resist an adversary able to request fresh challenges; holding nonce policy constant, in-proof context binding is the only strategy blocking same-epoch transfer under shared nonces. Measurements across six network conditions, seven venues in four countries, and an epoch-window simulation indicate same-epoch transfer is a realistic concern in dense urban deployments. Evaluation spans five platforms, seven strategies, and an end-to-end transfer attack; all artifacts are public.

cs.CR↗