arXiv · 2609.38005
Diagnosing and Improving Probabilistic Reasoning in Large Language Models
Abstract
Large language models (LLMs) are increasingly proposed as decision assistants who must reason probabilistically from available evidence under explicit decision costs. We propose a decision-theoretic framework that decomposes LLMs' decision loss into two components: forming accurate beliefs from provided evidence and translating those beliefs into actions that optimize a provided utility function. Using a synthetic benchmark with known ground truth, we apply the decomposition to characterize probabilistic reasoning in frontier and open-sourced models. We further evaluate whether RL interventions targeting beliefs, decisions, or both improve these components across three domains, whether improvements transfer across components and elicitation formats, and whether decision performance can improve without improvement in belief formation. We find that targeting one component of probabilistic reasoning redistributes decision loss, improving the target without necessarily transferring to others, and that jointly targeting belief formation and decision-making improves both but hinges on matched formats between training and evaluation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Huaman Sun, Dingcheng Wang, Jason Hartline, Jessica Hullman. 2026-09-29. Diagnosing and Improving Probabilistic Reasoning in Large Language Models. https://arxiv.org/abs/2609.38005
Cite the original work for its findings. Save a collection to share your selection of sources.