arXiv · 2609.26060
ChainUQ: Reasoning Consistency-Aware Uncertainty Quantification for Large Language Models
Abstract
While large language models (LLMs) exhibit impressive reasoning capabilities, response-level confidence may remain unreliable when intermediate claims conflict with the final conclusion. Therefore, effective uncertainty quantification (UQ) is required to capture logical inconsistencies within the reasoning chain, not just the correctness of the final output. Current approaches have two major limitations: (1) their reliance on token-level probabilities fails to capture reasoning consistency, and (2) they lack mechanisms to dynamically calibrate confidence using the structural logic of the generated chain. To advance existing research, we introduce ChainUQ, a reasoning consistency-aware uncertainty quantification framework for LLMs. ChainUQ consists of two key technical components: an alignment-aware lightweight UQ module that estimates a raw intrinsic model confidence score from frozen features aligned to the final conclusion, and a reasoning consistency-aware calibrator that refines this score using reasoning-chain consistency evidence. Evaluations across diverse in-distribution and out-of-distribution benchmarks show that ChainUQ consistently improves response-level uncertainty estimation, achieving an average 3.1% relative gain in AUROC and up to 45.0% relative reduction in ECE, and can be directly transferred to new settings without additional fine-tuning.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Dahai Yu, Rongchao Xu, Lin Jiang, Ximiao Li, Guang Wang. 2026-08-08. ChainUQ: Reasoning Consistency-Aware Uncertainty Quantification for Large Language Models. https://arxiv.org/abs/2609.26060
Cite the original work for its findings. Save a collection to share your selection of sources.