arXiv · 2608.20347
Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias
Abstract
Language models (LMs) often pass behavioral bias evaluations, but it remains unclear whether they no longer represent the underlying associations that give rise to biases, or have merely learned not to express them. In this study, we show that representational biases are often detectable, even when behavioral biases are not visible. We introduce a causal framework that decomposes occupational bias into two measurement points: a model's internal representation of a user's competence, and its observable outputs. We derive steering vectors for representations of user expertise, and verify that they causally mediate model behavior in both a question-answering task and a hiring task. Applying this framework to several open-weight models, we find that demographic attributes, such as gender, race, and socioeconomic status, influence a model's representation of user expertise, even in cases where behavioral metrics detect no disparity between demographics. We show that these model representations can influence downstream behavior under intervention, suggesting failure modes that behavioral metrics alone may not detect.
Explore related subjects
Keep this discovery
Keren Fuentes, Aaron Mueller. 2026-06-15. Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias. https://arxiv.org/abs/2608.20347
Cite the original work for its findings. Save a collection to share your selection of sources.