arXiv · 2609.32449
Self-Reports Do Not Identify Self-Models: An Identifiability Test for Counterfactual Reports
Abstract
Language-model self-reports are evidence about behavior in a prompt environment, not by themselves evidence of a self-model. We investigate counterfactual reports about affect-like states under activation interventions and ask whether the report remains bound to the named intervention when the demonstration environment changes. Across three open instruction models, wrong-source demonstrations move reports toward the source answer family, while explicit mechanism binding reduces this pull. Self-report benchmarks should include environment-shift invariance tests under fixed intervention before treating accuracy as evidence for an autonomous report mechanism.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Phongsakon Mark Konrad, Toygar Tanyel, Serkan Ayvaz. 2026-09-26. Self-Reports Do Not Identify Self-Models: An Identifiability Test for Counterfactual Reports. https://arxiv.org/abs/2609.32449
Cite the original work for its findings. Save a collection to share your selection of sources.