arXiv · 2509.24974
Double Descent as a Lens for Sample Efficiency in Autoregressive vs. Discrete Diffusion Models
Abstract
Data scarcity drives the need for more sample-efficient large language models. In this work, we use the double descent phenomenon to holistically compare the sample efficiency of discrete diffusion and autoregressive models. We show that discrete diffusion models require larger capacity and more training epochs to escape their underparameterized regime and reach the interpolation threshold. In the strongly overparameterized regime, both models exhibit similar behavior, with neither exhibiting a pronounced second descent in test loss across a large range of model sizes. Overall, our results indicate that autoregressive models are more sample-efficient on small-scale datasets, while discrete diffusion models only become competitive when given sufficient capacity and compute.
Explore related subjects
Keep this discovery
Ahmad Fraij, Sam Dauncey. 2025-09-29. Double Descent as a Lens for Sample Efficiency in Autoregressive vs. Discrete Diffusion Models. https://arxiv.org/abs/2509.24974
Cite the original work for its findings. Save a collection to share your selection of sources.