arXiv · 2606.03982
Language Models Compare Quantities Using Number-specific and Unit-specific Heuristics
Abstract
Quantities with measurement units, such as 110 cm and 1.2 m, require language models (LMs) to combine a numeral with a symbolic unit scale. Here, we study how LMs compare such quantities in controlled settings spanning several unit systems. We find that accuracy degrades near the comparison boundary, where small changes in value determine the correct answer. The resulting errors are systematic: linear surrogate models predict LM preferences from numerical-difference and unit-scale-difference cues, and causal interventions on subspaces aligned with these variables shift model's output. The results suggest that LMs compare quantities through a bag of heuristics over numerals and units, rather than first converting both expressions to an exact shared-scale representation.
Explore related subjects
Keep this discovery
Mutsumi Sasaki, Go kamoda, Ryosuke Takahashi, Kosuke Sato, Kentaro Inui, Keisuke Sakaguchi, Benjamin Heinzerling. 2026-06-02. Language Models Compare Quantities Using Number-specific and Unit-specific Heuristics. https://arxiv.org/abs/2606.03982
Cite the original work for its findings. Save a collection to share your selection of sources.