arXiv · 2607.27731
Towards joint scaling laws with optimal batch size schedules
Abstract
Modern deep learning typically keeps the batch size static throughout training, thus overlooking the joint effect of learning rate and batch size on the training dynamics. In this paper, we study the deep learning dynamics through the lens of convex optimization and derive a joint characterization of loss in terms of both schedules, applicable to general optimizers and model architectures. This characterization yields a closed-form optimal batch size schedule for any prescribed learning rate schedule, and further leads to joint scaling laws that consistently outperform static batch size baselines, highlighting the significance of dynamic batch size schedule in large language model training.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jiaxiang Li, Zhiqi Bu, Shiyun Xu. 2026-07-30. Towards joint scaling laws with optimal batch size schedules. https://arxiv.org/abs/2607.27731
Cite the original work for its findings. Save a collection to share your selection of sources.