arXiv · 2609.21096
Detecting Hallucination in LLMs: Tracing the Topological Signatures of Impaired Context Sharing
Abstract
In this work, we examine the topology of information flow patterns within attention graphs to effectively distinguish hallucinated from non-hallucinated responses. We analyze the Forman-Ricci curvature to identify structural patterns indicating information bottlenecks in attention graphs. We then introduce a method that captures both semi-local and global information-flow characteristics of attention heads associated with hallucinated responses. We evaluate our approach extensively across several LLMs and established benchmarks. Empirical results demonstrate that our proposed single-pass approach provides consistent improvements over existing attention-based and multi-response baselines across two hallucination-detection benchmarks, while achieving competitive performance across diverse LLM architectures. Further analysis reveals that impaired context sharing among tokens during causal generation is strongly associated with hallucination occurrences in LLMs. In particular, hallucinated responses are consistently characterized by an over-reliance on self-attention, diffused context retrieval from earlier tokens, or information over-squashing, especially in the final transformer layer.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Amir Jalilifard, Anderson Rocha, Eric Wong, Marcos Medeiros Raimundo. 2026-09-17. Detecting Hallucination in LLMs: Tracing the Topological Signatures of Impaired Context Sharing. https://arxiv.org/abs/2609.21096
Cite the original work for its findings. Save a collection to share your selection of sources.