arXiv · 2609.22522
When Cosine Similarity Fails to Reflect Linearly Accessible Structure in Dialogue Models
Abstract
Cosine similarity is widely used to analyze transformer representations, implicitly assuming that similarity reflects task-relevant structure. We study when this assumption fails in dialogue-conditioned large language models. Across three 7-8B chat-tuned models, ambient cosine similarity substantially underestimates linearly decodable persona structure on the same hidden states; numerically, linear probe AUC is in the 0.73-0.97 range while cosine kNN is in the 0.56-0.77 range on a 30-class task. A low-dimensional supervised subspace recovers much of this gap, whereas a matched-rank PCA subspace does not and in some cases degrades performance. This mismatch is regime-dependent: it is absent in single-sentence sentiment classification (SST-5), and a matched-cardinality control rules out attribute cardinality as a confound. The gap does not systematically increase across dialogue turns, and the task-aligned subspace remains stable over time. However, two of three models violate a pre-registered within-subspace separability invariance criterion (|Delta AUC| <= 0.03), and one model violates a pre-registered turn-invariance criterion (|Delta L| <= 0.05). These results show that cosine similarity can fail to reflect task-aligned structure in dialogue representations even when that structure is linearly accessible.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yu Sun, Mengyin Lu, Cong Feng, Guangming Lu, Huimin Han. 2026-09-18. When Cosine Similarity Fails to Reflect Linearly Accessible Structure in Dialogue Models. https://arxiv.org/abs/2609.22522
Cite the original work for its findings. Save a collection to share your selection of sources.