Searcharxiv⌕ Search

arXiv subjects

Richard Zhe Wang

Publications and source records attributed to Richard Zhe Wang.

3 recordsLinked to original sources

Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention

Softmax attention has two structural gaps. A head cannot abstain, because its weights sum to one, so it outputs something even when nothing is relevant. Nor can it filter what it reads, because its output is a weighted average of value vectors, passing interference as faithfully as signal. We call these missing primitives abstention and noise filtering. Recent studies report that gating the value pathway improves pretraining but attribute the gain to different causes. We show that a value gate partly supplies both primitives, which unifies the reported causes as views of one gain. We give each primitive its own mechanism in matched models of 10M to 350M parameters and measure what each contributes. The gain from gating is almost entirely abstention at 10M, whereas by 350M filtering contributes as much as abstention, so what a study observes depends on its scale. The two benefits are largely additive, with a small overlap. A gate determined by each value alone leaves the attention sink in place, whereas a query-controlled mechanism removes it. Injecting interference into the value reads shows that abstention and filtering protect against it in distinguishable ways. The same patterns appear in pretrained models up to 20B parameters.

cs.LG↗

The Communication Map of a Transformer

The components of a transformer communicate by writing to and reading from a shared residual stream, and the mechanistic interpretability literature has mapped these connections by hand, one circuit at a time. We present the communication map, which charts every potential communication channel from the geometry of the model's weights alone, generalizing the composition score of Elhage et al. (2021) into a single coupling coefficient covering all 18 connection classes, from head-to-head to neuron-to-neuron and everything in between. We provide an account of the properties of the coupling coefficient, including its geometric interpretation and its exact chance level. The census finds that 70-89% of head pairs are oriented far from chance, some coupled strongly and others actively avoiding each other. We demonstrate the communication map in two novel applications. In Application 1, we recover the known induction circuits blind from the strongest head-to-head couplings and group the heads into communities, and ablating one such community destroys the model's in-context copying. In Application 2, we pool the coupling coefficients of every head to identify a distinct two-dimensional residual stream subspace, whose deletion abolishes the induction capability in six models up to Pythia-6.9B. We show that this subspace is different from those identified by either activation PCA or outlier dimensions. We release the map, the statistical machinery, and the intervention suite.

cs.LG↗

Confidently Wrong: Detecting Hallucinations in Financial Question Answering from LLM Internal States

Large language models (LLMs) in financial applications fail most consequentially when they are confidently wrong. Hedged, uncertain answers invite scrutiny, whereas confident errors silently degrade downstream decisions without warning. We ask how reliably such confidently wrong answers, or confident hallucinations, can be detected from a model's internal activations, and whether those activations carry information beyond its observable outputs. We train linear probes on the residual stream and evaluate them on two established question-answering (QA) benchmarks built from real filings, FinQA and TAT-QA. Behavioral confidence is measured as the agreement among eight resampled answers to the same question, and probe effectiveness is compared against baselines, such as token log-probabilities and the model's own True/False self-assessment of its answer. Our findings show that among confident answers, those for which all eight resamples agree, 15-23% are wrong on FinQA. There the probes have a significant advantage over baseline methods in detecting hallucinations, holding 0.68-0.77 AUROC while the best baselines fall to 0.55-0.63, across Qwen3-8B, Llama-3.1-8B, and Gemma-2-9B. Our results suggest that probing can be a cost-effective triage mechanism for routing LLM answers to human review and quality control procedures in high-stakes financial applications.

cs.CL↗