arXiv · 2609.37737
Where Do LLMs Decide to Break the Rules? Mechanistic Localization of Prompt Injection Compliance
Abstract
When a prompt injection attack succeeds, a Large Language Model (LLM) abandons its assigned system role to comply with an adversarial instruction. While prior work has extensively quantified how often this occurs, we ask a more fundamental question: where inside the network does the model actually decide to break the rules? Using layer-by-layer causal activation patching across five models (4B to 32B parameters), we find a clear dissociation: attack information is linearly decodable from the first layer, yet causal leverage over the model's behavior is negligible until a late-layer bottleneck in the final third of the network. Patching this bottleneck reverses compliance in 77--92\% of cases. We show that the compliance mechanism occupies a compact linear subspace (rank-8 in 4B and 14B models, scaling to rank-64 at 32B) and is architecturally stable across varying model families. Finally, we validate our mechanistic account by showing that this causal peak layer is also the representationally optimal site for detecting attacks, outperforming early-layer classifiers that degrade under surface-level obfuscation such as leetspeak substitution. This alignment between causal leverage and detection performance provides converging evidence that the late-layer bottleneck captures decision-relevant computation rather than merely reflecting an artifact of the intervention.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Rui Wen, Jiayang Liu, Zeyu Yang, Jun Sakuma, Lu Sun. 2026-09-29. Where Do LLMs Decide to Break the Rules? Mechanistic Localization of Prompt Injection Compliance. https://arxiv.org/abs/2609.37737
Cite the original work for its findings. Save a collection to share your selection of sources.