arXiv · 2607.23130
Identification and Honest Recovery from Semantic Observation Kernels: Operator Error, Coarsening, and Stability
Abstract
Probabilistic text generators, such as large language models, assign probabilities to phrases, but consequential decisions require posterior uncertainty over meaningful states. These are not interchangeable: language probabilities depend on the prompt, may be incomplete and need not reliably identify state uncertainty. Without a statistical bridge, fluent responses and numerical confidence are insufficient for inference or governance. We formulate recovery of the target posterior as a semiparametric inverse problem and develop honest recovery guarantees that account jointly for calibration error, measurement noise, incomplete probabilities and weak identification. Simulations demonstrate the predicted coverage and stability behaviour, while two frozen language-model studies demonstrate held-out recovery. The resulting method determines when a semantic measurement can be trusted for inference and when use, review, recalibration or abstention is warranted, providing a statistical foundation for runtime AI governance.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Matthew Francis Dixon. 2026-07-25. Identification and Honest Recovery from Semantic Observation Kernels: Operator Error, Coarsening, and Stability. https://arxiv.org/abs/2607.23130
Cite the original work for its findings. Save a collection to share your selection of sources.