arXiv · 2609.33851
Rethinking Contextualization by Reinterpreting Attention Head Channels
Abstract
Contextualization, the core operation of language modeling, transmits information across words to build sentence-specific word representations. Prior works mainly study contextualization, focusing on individual words and attention heads as a growing discrete dictionary, lacking a global view of their general behavior. Therefore, we propose a general principle: Globally, we find and estimate that different words carry different amounts of information, and less-informative words tend to absorb more contextual information. Specifically, these low-information words do not absorb contextual words uniformly, and finer-grained selectivity enables more precise routing to promote information transmission between matched words. Moreover, to find what mechanism causes such processing, we reinterpret attention heads as channels gated by their singular vectors and find that: (1) these singular vectors point to the hidden states of more informative words, allowing such words to write their information to others more strongly to act as information sources, and vice versa; and (2) these singular vectors can be viewed equally as hidden state features, enabling automated interpretation of attention heads beyond prior heuristic head discovery, also embedding heads into a continuous space rather than treating them as discrete, independent dictionary entries.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Hakaze Cho, Haolin Yang, Zhun Sun, Naoya Inoue, Benjamin Heinzerling, Kentaro Inui. 2026-09-27. Rethinking Contextualization by Reinterpreting Attention Head Channels. https://arxiv.org/abs/2609.33851
Cite the original work for its findings. Save a collection to share your selection of sources.