arXiv · 2506.02285
Why Gradients Rapidly Increase Near the End of Training
Abstract
During long-duration Large Language Model (LLM) training runs the gradient norm increases rapidly near the end of training. In this short note, we show that this increase is due to an unintended interaction between weight decay, normalization layers, and the learning rate schedule. We propose a simple correction that fixes this behavior while also resulting in lower loss values throughout training.
Explore related subjects
Keep this discovery
Aaron Defazio. 2025-06-02. Why Gradients Rapidly Increase Near the End of Training. https://arxiv.org/abs/2506.02285
Cite the original work for its findings. Save a collection to share your selection of sources.