arXiv · 2510.03280
Training Optimal Large Diffusion Language Models
Abstract
We introduce Quokka, the first systematic scaling law for diffusion language models (DLMs), encompassing both compute-constrained and data-constrained regimes, and studying the key modeling and optimization designs. Quokka is a good friend of Chinchilla and provides wider scopes. We hope the results would bring short-term practical guidance in DLMs training and long-term inspirations for the whole AI community.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jinjie Ni, Qian Liu, Chao Du, Longxu Dou, Hang Yan, Zili Wang, Tianyu Pang, Michael Qizhe Shieh. 2025-09-28. Training Optimal Large Diffusion Language Models. https://arxiv.org/abs/2510.03280
Cite the original work for its findings. Save a collection to share your selection of sources.