arXiv · 2507.22748
Nine Raters, One Index: Carrying LLM Disagreement into Labour-Market Estimates
Abstract
When a large language model supplies research annotations, the choice of model becomes an analytic degree of freedom. We ask how much that choice matters by putting 1,100 task-occupation cells to nine models from three vendors and weighting their ratings with worker-reported task importance in the British Skills and Employment Surveys. Rankings are substantially more stable than levels: pairwise rank correlations range from 0.74 to 0.92, while the share of British jobs scoring above 0.5 ranges from under 0.1% to 38%. We therefore treat the rater as one dimension of a multiverse, reporting each rater-specific estimate and its range alongside a Rubin-style summary that incorporates their observed dispersion. The consequences depend on the downstream design. A post-2022 pay gradient of -0.020 log points per interquartile range is negative under all nine raters, although it continues a pre-existing trend. In online vacancies, all nine raters initially produce negative post-2022 gradients; after teleworkability is given its own period path, seven produce significant positive coefficients, two produce small negative coefficients, and the pooled estimate is indistinguishable from zero. Neither application identifies an effect of generative AI. Validation against worker-reported AI use, competing exposure measures, assistant traffic and an independent task survey supports incremental predictive content and portability, but any bias shared across raters remains unidentified. The paper's contribution is a procedure for exposing and carrying observed LLM-rater disagreement into empirical estimates.
Explore related subjects
Keep this discovery
Golo Henseke, Rhys Davies, Alan Felstead, Duncan Gallie, Francis Green, Ying Zhou. 2025-07-30. Nine Raters, One Index: Carrying LLM Disagreement into Labour-Market Estimates. https://arxiv.org/abs/2507.22748
Cite the original work for its findings. Save a collection to share your selection of sources.