arXiv · 2605.19095
ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models
Abstract
Schedule-Free Learning has shown promise as a practical anytime training method for machine learning, showing success across dozens of standard benchmark problems. However, strong performance for LLM training has only been demonstrated at small scales. We identify a number of fixes necessary to scale up Schedule-Free Learning to larger batch sizes and model sizes, and present a learning-rate-free and schedule-free method (ScheduleFree+) for training large language models which greatly outperforms Warmup-Stable-Decay (WSD) schedules. We also demonstrate that Schedule-Free Learning is most effective for long duration training, and at 1000 tokens per parameter, it outperforms SOTA schedules by 31%. Schedule-Free Learning provides a theoretical foundation for the use of model averaging and checkpoint merging during pretraining.
Explore related subjects
Keep this discovery
Aaron Defazio. 2026-05-18. ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models. https://arxiv.org/abs/2605.19095
Cite the original work for its findings. Save a collection to share your selection of sources.