TY - RPRT TI - Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful AU - Martin Marek AU - Sanae Lotfi AU - Aditya Somasundaram AU - Andrew Gordon Wilson AU - Micah Goldblum PY - 2025 UR - https://arxiv.org/abs/2507.07101 ID - 2507.07101 ER -