A Solvable Theory of Pre-training Data Poisoning: Regime-Dependent Scaling Exponents
Pre-training data poisoning of large language models is usually studied using targeted backdoors and their survival through safety post-training, which leaves open a more basic question: how does a model's clean data performance degrade as the poison rate $\varepsilon$ grows? Motivated by our controlled pre-training runs of OLMo-style models, in which the relative clean data validation perplexity increase $Δ$ between poisoned and clean models matched in architecture, token budget, and optimization schedule is well fit by a power law $Δ\approx C\varepsilon^{a}$ with a non-integer exponent, we ask what such a law requires theoretically. We first prove an analyticity barrier: whenever the contaminated objective depends analytically on $\varepsilon$ around a nondegenerate clean data optimum, $Δ$ is generically quadratic in $\varepsilon$, so a generic non-integer exponent is a signature of genuinely singular structure. We then supply that structure in solvable truncated ridge regression with heavy-tailed covariates, controlled by $q_\star$, and a label-shift poisoning. Our central result is that the excess risk scaling exponent depending on the order of limits: in the higher dimensional proportional regime it is $ε^{q_\star/(q_\star+2)}$, whereas taking the ample-data limit first gives $ε^{2-2/q_\star}$, and the limits do not commute. We confirm this prediction through several numerical simulations. Finally, we argue that finite training time plays the role of truncation on the curvature spectrum in local LLM pre-training, deriving the observed scaling law under heavy tailed inverse curvature spectrum as a modeling hypothesis.