arXiv · 2602.09591
On the Optimal Reasoning Length for RL-Trained Language Models
Abstract
Reinforcement learning substantially improves reasoning in large language models, but it also tends to lengthen chain-of-thought outputs and increase computational cost. Although length-control methods have been proposed, the length-accuracy relationship they induce remains unclear. We train policies with several length-control methods on multiple base models in a controlled setup and find that, across both mathematical reasoning and code generation, accuracy is non-monotonic in output length, peaking at an intermediate value. Mode accuracy, however, continues to improve with length even in settings where sample accuracy plateaus or declines, indicating that the non-monotonic length-accuracy relationship is driven by dispersion around an increasingly correct center.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Daisuke Nohara, Taishi Nakamura, Rio Yokota. 2026-02-10. On the Optimal Reasoning Length for RL-Trained Language Models. https://arxiv.org/abs/2602.09591
Cite the original work for its findings. Save a collection to share your selection of sources.