arXiv · 2609.16648
GrowMTP: Can RL Grow Its Own Draft Head?
Abstract
Reinforcement learning (RL) post-training drives the frontier capabilities of large language models, with its wall-clock dominated by autoregressive rollout generation. Speculative decoding is an established remedy for this bottleneck, but existing draft heads must be pretrained or warmed up before RL, introducing substantial training cost outside the RL run to be accelerated. We observe that RL training itself provides both conditions required for online draft-head training: its rollout distribution is far narrower than that of pretraining, and its verification step continuously produces supervision signals aligned with this distribution. Building on these observations, we propose GrowMTP, which uses this supervision to train a draft head from scratch entirely within the RL loop, with all head updates detached from the policy backbone. On Qwen3-4B (no draft head), MiMo-7B-SFT (weak head), and Qwen3.5-4B-Base (strong head), GrowMTP achieves rollout speedups of 2.13x, 1.93x, and 1.36x, and end-to-end speedups of 1.60x, 1.41x, and 1.20x, respectively. GrowMTP therefore serves existing RL training frameworks as a modular component, particularly offering a from-scratch acceleration path for models without pretrained draft heads.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Minghua He, Lingzhe Zhang, Yuan Liu, Xiao Zhou, Aiwei Liu. 2026-09-15. GrowMTP: Can RL Grow Its Own Draft Head?. https://arxiv.org/abs/2609.16648
Cite the original work for its findings. Save a collection to share your selection of sources.