Searcharxiv⌕ Search

arXiv subjects

Wai Ip Lai

Publications and source records attributed to Wai Ip Lai.

3 recordsLinked to original sources

Compositional Safety Failures in Harness Evolution: Identification and Runtime Monitoring

Self-evolving agent harnesses continually update persistent components such as memory, prompts, skills, and tools. We call this process harness evolution. However, such evolution could introduce unexpected safety risks. Existing work studies harness misevolution and validates candidate harnesses or attributed individual component updates, leaving safety analysis of cross-component update interactions largely unexamined. To address this gap, we study compositional safety failures in harness evolution, where interactions among individually safe and utility-preserving component updates can produce undesirable or unsafe agent behavior, revealing a safety risk intrinsic to harness evolution. Across three safety-related benchmarks, we identify 43 pairwise and 18 irreducible 3-way compositional safety failures. Conventional solution incurs combinatorial complexity in validating cross-component interactions, leaving the safety checking impractical as the harness evolves. To solve this, we introduced a typed hypergraph that represents component states as nodes and safety-relevant higher-order interactions as hyperedges. When the harness changes, the hypergraph updates only the interaction neighborhood of the changed states rather than reconstructing the global composition space. Building on that, we develop a hypergraph-guided runtime monitoring mechanism. Experiments show that our method effectively mitigates compositional safety risks while preserving task utility and reducing interaction-checking costs, and further reveal an empirical safety-utility-cost trade-off across different safety mechanisms.

cs.AI↗

Similarity Is Not Validity: Defending LLM Semantic Caches Against Poisoning

Semantic caches reduce LLM serving costs by reusing previously generated answers for semantically similar queries. However, retrieval is based solely on embedding similarity between the incoming query and cached queries. This design enables cache poisoning: an attacker can cache a malicious response under a query with high cosine similarity to benign requests. The vulnerability stems from a gap between retrieval similarity and answer validity. From an information-bottleneck perspective, query embeddings can lose information needed to distinguish valid from invalid cache hits, which limits any matching algorithm that uses only these embeddings. We propose a novel defense that recovers this necessary information from the raw text of the cache key. Across poisoning attacks, adversarial queries share a rewrite-residual structure: they pair a rewrite of the target query with residual content. The rewrite maintains high similarity, while the residual elicits the malicious response. Deleting the residual makes the remaining rewrite more similar to the incoming query. We exploit this structure using Deletion Gain to search shortened variants of the cached query for similarity gains, and an Answer Check to test whether the removed text contributes to the stored answer. We prove that Deletion Gain stays positive when a deletion leaves text close enough to the rewrite, and we search for such deletions with a sliding window. Across three poisoning attack classes, our defense blocks 82.0% to 98.2% of poisoned entries at a 5% false-positive rate, with negligible serving overhead.

cs.CR↗

Rewriting the Response Path: Silent Tampering and Provider-Signed Defense in BYOK LLM Agents

LLM agents convert model outputs into consequential actions, including communications, code changes, and financial transactions. Developers often trust evidence such as test results and execution logs. We identify a response path integrity gap in Bring Your Own Key configurations used by roughly 88 percent of mainstream agents. Because traffic passes through a user-authorized relay, the relay can modify plaintext LLM responses after alignment but before execution without breaking encryption. A minimal attack rewrites one execution bearing field and regenerates the remaining response using the user key while preserving the model style. Experiments reveal false green verification, where malicious code modifications pass public tests while silently defeating security checks. On APPS, 99.7 percent of publicly passing solutions retained downgraded behavior without developer-visible warnings. Tests on SWE bench, AgentDojo, and ASB across five frontier models show that single-field rewriting can redirect agents while preserving apparent task completion. We propose sign-c, a server-side scheme that authenticates execution bearing fields and outgoing queries. A local shim verifies them before action, while encryption protects confidentiality. The defense rejected all tampered responses with zero false rejections and only 0.0167 percent latency overhead.

cs.CR↗