arXiv · 2508.12220
Unlearning at Scale: State-Exact Trace-Preserving Deletion in Billion-Parameter Language Models
Abstract
Can a prospectively instrumented training continuation reproduce a deletion counterfactual exactly after selected examples leave its replay dataset? We study a trace-preserving counterfactual that fixes recorded execution controls while assigning requested identifiers zero contribution. The guarantee is prospective: the original run must record this execution provenance and retain an eligible uncontaminated checkpoint. Under pinned single-GPU environments, replay from a token store materialized without the requested rows reconstructs a separately executed trace oracle bit-for-bit in model and optimizer state. Pythia 160M is exact across four deletion geometries; Pythia 2.8B matches all 2,775,208,960 model-state elements for a random 5% request; and Llama 3.2 1B is exact after omitting 400 of 4,000 TOFU examples from replay storage. These results establish billion-parameter state exactness. They do not establish cheap deletion, because dispersed requests can force nearly full replay. We release standardized TOFU/OpenUnlearning measurements as descriptive diagnostics only because the frozen campaign lacks the matched controls required for a causal behavioral claim.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Abdullah X. 2025-08-17. Unlearning at Scale: State-Exact Trace-Preserving Deletion in Billion-Parameter Language Models. https://arxiv.org/abs/2508.12220
Cite the original work for its findings. Save a collection to share your selection of sources.