arXiv · 2609.38011
When do data mixtures improve scaling laws? Insights from high-dimensional regression
Abstract
Modern machine learning systems are trained on mixtures of data from different domains, and choosing the right mixture can substantially improve downstream performance. Despite an extensive literature on data mixing and reweighting, existing work is largely empirical and it remains unclear when auxiliary data genuinely improves scaling laws rather than merely providing more samples. To gain insight into this question, we study a high-dimensional mixed-data regression model with a shared regression function, heterogeneous covariances and noise levels, and dataset sizes that may grow at different rates. We establish the minimax risk under an ellipsoidal parameter constraint for the general covariance structure and derive deterministic equivalents for the test error of ridge regression under commutative covariances. We then specialize to a target domain and an auxiliary domain with aligned power-law covariance spectra, where the theory yields explicit scaling laws in terms of spectral decay, target regularity, and the relative growth of the two datasets. These laws identify regimes in which combining data mixtures provably yields a faster scaling rate than using either dataset alone. In particular, improving the scaling law requires a specific interplay between spectra and relative sample sizes of the domains. Our numerical experiments on language models exhibit the same qualitative phenomenon: appropriate data mixtures yield a faster decrease in target-domain test loss than training on either domain alone.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Diyuan Wu, Lehan Chen, Theodor Misiakiewicz, Marco Mondelli. 2026-09-29. When do data mixtures improve scaling laws? Insights from high-dimensional regression. https://arxiv.org/abs/2609.38011
Cite the original work for its findings. Save a collection to share your selection of sources.