SearcharxivSearch

arXiv subjects

Sehee Hong

Publications and source records attributed to Sehee Hong.

2 recordsLinked to original sources

Reference-Distribution Dependence in LLM-Based Synthetic Persona Data: Diagnosis and Post Hoc Adjustment of Demographic Distributions

We diagnose how closely the demographic distributions in LLM-based synthetic persona data match external reference distributions. For the three variables examined, we show that most of the observed error is attributable to the choice of reference rather than to the generator. Using total variation distance (TVD), we compare the sex x age group x province joint distribution of 1,000,000 records from Nemotron-Personas-Korea (NPK) with Korean official statistics. Against resident-registration figures for April 2026, the time of use, the bias bound, defined as the largest possible difference in the share of any subgroup formed from the three variables, is 1.81 percentage points. This is comparable to the margin of error of a survey of roughly 2,900 respondents. This value is not a fixed property of the data. Matching the reference period and series to the generating reference identified here, the 2024 register-based census restricted to Korean nationals, lowers it to 0.56 percentage points. Over the 15 months between the best-matching month (January 2025) and the time of use, the resident-registration population structure itself moves more than twice the distance of NPK's minimum error. Raking and cell post-stratification, the two weighting schemes used in Korean survey practice, remove most of the reference-period dependence at a variance inflation of about 0.2% in both cases. After raking against the generating reference, the residual joint discrepancy lies at, and marginally above, the upper bound of what a perfect generator would produce when realizing 1,000,000 records (97.6th percentile of the Monte Carlo distribution). We recommend treating synthetic persona data as auxiliary material for small-scale survey design rather than as a substitute for survey data, and re-running both diagnosis and adjustment against official statistics current at the time of use.

cs.CY

Recovering Nonlinear Functions of Latent Variables: A Plausible-Value Neural Network Framework

When factor scores replace true latent scores in nonlinear prediction, measurement error attenuates the recoverable variance of any $k$th-order component of the regression function by $\rho^k$ -- the $k$th power of the score's coefficient of determination -- for any linear score type. This study derives the bound via Hermite polynomial expansion and proposes PV-ANN -- plausible values (posterior draws preserving latent variance) combined with artificial neural networks (learning functional form without prespecification). The bound governs recovery of the latent-scale function, not prediction of the outcome from observed indicators, for which factor scores are already sufficient; the two metrics are therefore predicted to dissociate. An 18-condition simulation supports both predictions: in the nonlinear low-reliability conditions PV-ANN closes about four fifths of the function-shape recovery gap between a factor-score learner and one given the true latent values, and the margin widens as reliability falls, while predictive accuracy is not improved, as the theory requires. A Big Five application illustrates the intended exploratory workflow and delineates boundary conditions under weak signal and measurement model misspecification.

stat.ME