arXiv · 2609.34828
Simulating Respondents, Not Single Questions: Coherent Survey Generation with Large Language Models
Abstract
Large language models are increasingly used to simulate response distributions in social surveys. Prior work has achieved accurate population-level simulation for individual questions. Real questionnaires, however, ask each respondent a sequence of related questions. A simulated respondent should show coherent preferences across the whole questionnaire, not merely accurate distributions for isolated items. Existing single-item methods cannot accurately reproduce how the same person answers a complete survey. We propose FullRespondent-LLM (FR-LLM), which fine-tunes two specialized LLMs: a marginal model for each item's response distribution and a respondent-level autoregressive model for dependencies across answers. Marginal-Constrained Joint Projection (MCJP) then projects the autoregressive joint distribution onto the set satisfying the item-level marginals learned by the first model. This yields complete questionnaires with realistic cross-item relationships while retaining strong item-level accuracy. On two real-world social survey datasets, FR-LLM more accurately reproduces multi-question response patterns, maintains competitive single-item accuracy, and generalizes better to unseen populations and questions. In a small commercial-survey dataset, we use simulated responses to make pricing and stocking decisions; FR-LLM achieves the highest realized profit.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ji Huang, Mengfei Li, Shuai Shao. 2026-09-28. Simulating Respondents, Not Single Questions: Coherent Survey Generation with Large Language Models. https://arxiv.org/abs/2609.34828
Cite the original work for its findings. Save a collection to share your selection of sources.