arXiv · 2609.34158
Toward a Graded Measure of Belief Stability in Large Language Models
Abstract
Large language models (LLMs) increasingly mediate how people access and reason with information, yet factual reliability is usually evaluated one judgment at a time. We introduce graded belief stability, a relational measure of how well a belief persists within an LLM's broader belief system. Unlike individual belief probability, it asks whether support for a claim persists when that claim is considered alongside the model's other epistemic commitments. We operationalize this idea with a Direct Conditional estimator that uses internal model representations to estimate conditional belief probabilities. Across 12 LLMs and three domains, lower-stability beliefs exhibit greater mean behavioral movement under conversational challenge in 83.3% of model-domain settings after matching on individual belief probability. Graded belief stability therefore extends reliability assessment beyond how strongly an LLM supports a claim to how robustly that belief is supported within its broader system of beliefs.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Samantha Dies, Branden Fitelson, Tina Eliassi-Rad. 2026-09-28. Toward a Graded Measure of Belief Stability in Large Language Models. https://arxiv.org/abs/2609.34158
Cite the original work for its findings. Save a collection to share your selection of sources.