arXiv · 2609.29845
Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs
Abstract
While Large Language Models (LLMs) rely on highly non-linear components, in this work we demonstrate that they exhibit fundamental linearity: when inputs from distinct text streams are linearly combined, the model outputs a superposition of the individual next-token distributions. We term this the \textit{Superposition Linearity Hypothesis}. We provide evidence that superposition is an intrinsic property of the Transformer architecture rather than an emergent consequence of training; in fact, we observe that it tends to diminish as pretraining progresses. However, we demonstrate that linearity can be substantially restored through lightweight fine-tuning, significantly reducing the divergence between the predicted next-token distribution and the average of the individual next-token distributions. Finally, we introduce a guided decoding procedure that disentangles superposed outputs, enabling the simultaneous generation of two coherent continuations from a single forward pass.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Pavel Tikhonov, Anton Korznikov, Matvey Mikhalchuk, Nikita Dragunov, Temurbek Rahmatullaev, Polina Druzhinina, Anton Razzhigaev, Ivan Oseledets, Elena Tutubalina. 2026-09-24. Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs. https://arxiv.org/abs/2609.29845
Cite the original work for its findings. Save a collection to share your selection of sources.