SearcharxivSearch

arXiv subjects

M. Moran

Publications and source records attributed to M. Moran.

2 recordsLinked to original sources

LLMs Show No Signs Of Individuated Metacognition

Confidence-weighted routing, selective abstention, and ensemble weighting all assume that a model's stated confidence is informative about its capability on the question being asked. They presume functional metacognition, the capacity to assess one's own capabilities, without exercising them. Aggregate calibration is well studied, with mixed results, but the underlying structure of elicited confidence is less well understood. We decompose binary confidence judgements from 20 frontier Large Language Models (LLMs) across six benchmarks using tetrachoric factor analysis paired with pairwise calibration, asking whether two models that differ in confidence also differ in performance. On factual recall and information retrieval benchmarks the cross-model confidence matrix is approximately rank-one and a single dominant factor captures most of the latent variance. Models retrieving facts share an item-level difficulty axis and differ mainly in their decision thresholds along it. Across all benchmarks the relationship between confidence and performance collapses once items that all models agree on are removed. Inter-model pairwise calibration is small even where statistically significant, and what remains shrinks to nothing once base-rate differences along the shared factor are controlled for. Mathematical reasoning is the apparent exception, but this turns out to be a confound where reasoning models answer questions about their confidence by trying to solve them in their chain of thought, bypassing the sub-symbolic self-knowledge we seek to measure. We find no evidence for significant verbalised individuated metacognition in any tested domain.

cs.LG

Measurement of the $^{58}$Ni($α$,$γ$)$^{62}$Zn reaction and its astrophysical impact

Cross section measurements of the $^{58}$Ni($α$,$γ$)$^{62}$Zn reaction were performed in the energy range $E_α=5.5-9.5$ MeV at the Nuclear Science Laboratory of the University of Notre Dame, using the NSCL Summing NaI(Tl) detector and the $γ$-summing technique. The measurements are compared to predictions in the statistical Hauser-Feshbach model of nuclear reactions using the SMARAGD code. It is found that the energy dependence of the cross section is reproduced well but the absolute value is overestimated by the prediction. This can be remedied by rescaling the $α$ width by a factor of 0.45. Stellar reactivities were calculated with the rescaled $α$ width and their impact on nucleosynthesis in type Ia supernovae has been studied. It is found that the resulting abundances change by up to 5\% when using the new reactivities.

nucl-ex