arXiv · 2601.21624
Training Memory in Deep Neural Networks: Mechanisms, Evidence, and Measurement Gaps
Abstract
Modern deep-learning training is not memoryless. Updates depend on optimizer moments and averaging, data-order policies (random reshuffling vs with-replacement, staged augmentations and replay), the nonconvex path, and auxiliary state (teacher EMA/SWA, contrastive queues, BatchNorm statistics). This survey organizes mechanisms by source, lifetime, and visibility. It introduces seed-paired, function-space causal estimands; portable perturbation primitives (carry/reset of momentum/Adam/EMA/BN, order-window swaps, queue/teacher tweaks); and a reporting checklist with audit artifacts (order hashes, buffer/BN checksums, RNG contracts). The conclusion is a protocol for portable, causal, uncertainty-aware measurement that attributes how much training history matters across models, data, and regimes.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Vasileios Sevetlidis, George Pavlidis. 2026-01-29. Training Memory in Deep Neural Networks: Mechanisms, Evidence, and Measurement Gaps. https://arxiv.org/abs/2601.21624
Cite the original work for its findings. Save a collection to share your selection of sources.