arXiv · 2608.23005
Large language models simulate intersectional synthetic identities with a budget of one to two dimensions
Abstract
Large language models are increasingly used as synthetic survey respondents, promising cheap access to rare intersectional populations. We test standard demographic-persona methods against every real intersectional subgroup across 15 waves of Pew's American Trends Panel -- 21 million simulated response distributions from eight models. In real respondents, subgroup opinion is approximately the additive sum of its single-identity components, yet grows 2.5x more distinctive as identities intersect. Simulated respondents show no such composition: a single feature explains a two-feature persona's responses better than the additive combination in 75-82% of subgroups, and a third feature adds almost nothing. This collapse survives every prompting strategy we test. Additionally, the feature models retain is chosen nearly blindly -- except that they systematically discard race and religion, the strongest real drivers of opinion. Synthetic samples offer intersectional personas but represent one identity at a time.
Explore related subjects
Keep this discovery
Virgile Rennard, Christos Xypolopoulos. 2026-08-24. Large language models simulate intersectional synthetic identities with a budget of one to two dimensions. https://arxiv.org/abs/2608.23005
Cite the original work for its findings. Save a collection to share your selection of sources.