arXiv · 2609.37700
Locating Answer-Correctness Signals in Frozen Large Language Models
Abstract
Language models expose internal signals that predict whether an answer is correct, readable from a single forward pass of a frozen model without additional generations. Yet existing probes often commit to one signal family or layer and can be brittle under distribution shift; in retrieval-augmented settings, many specialized detectors instead target passage faithfulness, which can diverge from correctness when retrieved evidence is unhelpful or conflicting. We therefore ask where answer correctness is readable, which internal signal families carry it, and how they should be combined. We search over hidden states, token probabilities, residual-stream features, attention, and their fusion, treating the selected readouts as a predictive measurement rather than a mechanistic localization. We run this analysis separately in closed-book and with-context settings, since context can change which readouts are informative. A consistent anatomy emerges: correctness concentrates in the answer span, recovered from the answer tokens even under retrieval, and the families carry it complementarily, so fusing them helps most out of distribution, where a single signal is weakest. The protocol is effective across two backbones and gates a retrieval controller as one downstream use.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yuansen Liu, Yixuan Tang, Anthony Kum Hoe Tung. 2026-09-29. Locating Answer-Correctness Signals in Frozen Large Language Models. https://arxiv.org/abs/2609.37700
Cite the original work for its findings. Save a collection to share your selection of sources.