arXiv · 2609.35108
A mechanistic study of language model introspection
Abstract
Large language models (LLMs) can sometimes report perturbations to their internal activations---even when the input provides no evidence that an intervention occurred. How do models detect and localize such internal changes? We study this question using a controlled task that keeps the input text fixed. We either inject a concept vector into the hidden state at one of ten token positions or apply no intervention. The model is asked to identify the perturbed position or report that no intervention occurred. Across three model families, we identify two small groups of attention heads with distinct roles in introspective reporting. Middle-layer gate heads influence whether the model reports a change, while router heads in a later layer help select the position to report. Interventions on gate heads can suppress position reports even when router heads supply location information. We further examine why reporting accuracy varies across concepts. Concept vectors that are localized more accurately produce stronger attention-score and output responses in gate heads, which is associated with better alignment of the induced key and value changes in their QK and OV computations. Together, these findings identify attention-head mechanisms supporting introspective detection and localization.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jiahong Zou, Xiangkun Sun, Lingkai Kong, Tonghan Wang. 2026-09-28. A mechanistic study of language model introspection. https://arxiv.org/abs/2609.35108
Cite the original work for its findings. Save a collection to share your selection of sources.