arXiv · 2609.22206
Dissecting Training-Free Uncertainty Estimation in Multimodal Large Language Models
Abstract
Multimodal Large Language Models (MLLMs) have achieved remarkable performance across a wide range of multimodal tasks, yet understanding and quantifying their predictive uncertainty remains underexplored despite being central for safety critical applications. In this work, we present a systematic study of training-free uncertainty quantification strategies for MLLMs, categorizing existing approaches into three conceptual families: token-level methods, which operate directly in the text output space; verbalized methods, which elicit uncertainty estimates or abstention signals via natural language prompts; and semantic methods, which measure uncertainty in a semantic meaning space. We benchmark these strategies across multiple datasets, model families, generations, and scales, and find that no single family dominates: token-level entropy (at sampling temperature 1.0) wins on short answers, verbalized abstention on sentence-length responses, and semantic methods on long-form generation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Soroush Seifi, Vaggelis Dorovatas, Lin Li, Yarin Gal, Rahaf Aljundi. 2026-09-01. Dissecting Training-Free Uncertainty Estimation in Multimodal Large Language Models. https://arxiv.org/abs/2609.22206
Cite the original work for its findings. Save a collection to share your selection of sources.