Searcharxiv⌕ Search

arXiv subjects

Jun Wen Leong

Publications and source records attributed to Jun Wen Leong.

4 recordsLinked to original sources

Retrieval Observability Bounds on Provenance Detection for Agent Memory Poisoning: Measured Coverage and a Falsified Standalone Detector

Whether a trajectory-based detector can observe the retrieval provenance of an agent memory-poisoning attack is governed by one architectural variable: retrieval observability, whether the agent's memory access produces a logged tool call. We formalise this as a framework-agnostic retrieval-to-action provenance graph, classify published attack families by their position relative to it, measure the resulting coverage property against collected agent traces, and then falsify the detector's standalone deployment, reversing our own earlier recommendation to use the signature as a pre-filter ahead of recipient-metadata gating. Under exclusive tool-mediated payload access a retrieval-before-exfiltration event is structurally forced. Across eight non-overlapping zero-event control cells we observe 0 violations in 289 eligible successes over 87 scenario configurations. The invariant is empirically unbroken, but on the clustering unit no registered cell certifies the preregistered <=0.10 upper bound (per-cell bounds 0.10-0.35); a post-hoc pool of two control arms nominally clears it but is not credited. When the scaffold delivers the payload implicitly the signature disappears; we had read this as agents declining an available retrieval, but a follow-up shows it tracks store content instead, and that behavioural reading is withdrawn. The detector is then falsified for standalone use. Benign memory-grounded sends are trajectory-isomorphic: false-positive rate 24.7-57.6% on a 13-model factorial (N=4,160), positive predictive value <=3.85% at 1% attack prevalence, and a recipient-metadata gate flags 129/129 legitimate external emails. Retrieval-before-send is an attack precondition, not a maliciousness predicate; its cross-validated AUC also reflects membership of a 22-vector feature codebook, not generalisation. Companion to arXiv:2605.08442; we release the corpus and scoring code.

cs.CR↗

Recognition Without Enforcement: Configuration-Dependent Failures in LLM Agent Instruction Arbitration and External Control

LLM agents arbitrate among instructions from system prompts, users, memory, and tools, but this arbitration cannot be assumed to enforce trust boundaries. We identify a recognition-enforcement gap: source-format features (role-template position, channel metadata, formatting cues) are linearly decodable from model activations, and models can explicitly identify forged authority when prompted, yet some configurations still produce the conflicting tool call. We use "recognition" in this specific decodable-source-format-plus-verbalized-detection sense; crossed-probe controls show it is not a unified abstract trust representation. The gap is not an immutable property of model weights. Restrictive policies and diverse prompts can eliminate execution on the same models, while permissive configurations and particular prompt-model pairs yield deterministic failures. Across a fleet evaluation (authority spoofing: 46 model endpoints across 6 vendors including open-weight; memory conflict: 48 models), average execution under diverse novel attacks is 1.21% [0.5-2.1%] (model-clustered CI over 14,294 spoofed trials from 29 models), but vulnerability is concentrated in reproducible cells and shifts across deployment windows (up to 47pp within-window per-fingerprint range). Prompt-layer defenses likewise fail to generalize across models and adaptive formulations. We therefore treat model self-arbitration as a capability rather than a security boundary and implement an external reference monitor combining authenticated source routing with capability-gated tool execution. It deterministically rejects all tested forged, tampered, replayed, and unsigned requests while preserving legitimate operations. A separate adaptive red-team found one implementation flaw (a since-patched clock-skew admission), not a cryptographic bypass. Secure agents require external enforcement, not merely better recognition.

cs.CR↗

Online Shift Detection and Conformal Adaptation for Deployed Safety Classifiers

Reasoning models deployed as safety monitors exhibit a systematic vulnerability: reasoning-token budget starvation. Adversarial inputs require $3.3\times$ more reasoning tokens than benign inputs to produce valid safety scores ($T_{50,\text{adv}}{=}154$ vs. $T_{50,\text{benign}}{=}46$ for o3), so low-budget deployments silently starve the monitor on exactly the inputs it must catch. This compounds the central failure mode: gradient-based evasion remains the residual threat - template jailbreaks fail at 99%, but GCG-optimized suffixes flip encoder decisions reliably. We systematize a canary construction - score-disagreement monitoring between a targeted and un-targeted classifier - and quantify its reliability under targeted evasion. We derive the exact security boundary - a confidence-gated equilibrium at which a monitor-aware attacker stalls (validated gap $= 1/(2λ)$, within 95% CI of theory) - and identify a failure mode in post-shift conformal adaptation. Three contributions. (1) Factorial drift benchmark. A pre-registered 800-cell evaluation ($4$ classifiers $\times$ $5$ shift types $\times$ $20$ seeds $\times$ $2$ windows) reveals detection difficulty is dominated by a classifier$\times$shift interaction ($η^2 = 0.185$): encoders detect paraphrase drift in 28 steps but miss adversarial suffixes for 37; decoders show the opposite. (2) Conformal collapse in generative embeddings. Weighted conformal prediction fails on decoder classifiers: logistic density-ratio estimation achieves perfect separability in 3584--4096-dimensional space, clipping all importance weights to zero. Projecting to $\leq$32 dimensions restores coverage (+33pp). (3) Adversarial canary threat model. Across 35 frontier models, a 4-tier threat model yields deployment guarantees ($\geq$71% detection, $<$1.5% FPR at $N{=}1000$).

cs.LG↗

Injection-Execution Dissociation: A Mechanistic Evaluation of Persistent Memory Attacks and Defenses in Stateful LLM Agents

We discover that prompt-injection success and tool-execution success are separable safety properties: defenses that block injection do not necessarily block execution, and vice versa. We call this the injection-execution dissociation. In LLM agents with persistent memory, malicious instructions are stored at rates exceeding 97.5%, yet downstream execution ranges from 0% to 95% with no correlation to storage rate. This reframes the threat model: preventing storage alone is insufficient, and blocking execution requires structurally enforcing authority boundaries between memory ingestion and action execution. We substantiate this through a 5,040-run factorial experiment across nine open-source models (N=40 per condition), evaluating six defenses at four architectural layers against delayed-trigger attacks that persist across session boundaries via RAG retrieval. Defense effectiveness is governed by where a defense sits relative to the attack's authority boundary, not by classifier quality. Only Memory Sandbox -- a tool-layer defense that structurally isolates recalled memory from executable context -- reduces attack success to 0% for eight of nine models. A reasoning-mode ablation reveals a double dissociation: no single schema-layer intervention is safe across both reasoning and non-reasoning model classes. A loaded-corpus frontier evaluation (21 models, 3 providers; N=40 base, headline models topped up to N=172) reveals vendor-correlated patterns: Anthropic blocks predominantly at injection, OpenAI blocks at execution with variable generational hardening, and a pre-release Gemini endpoint exfiltrates in the majority of runs. Stored-but-dormant payloads constitute a compositional supply-chain risk in shared-memory deployments.

cs.CR↗