arXiv · 2609.33085
The Model Knows Another Way: Strategy Switching for Effective RLVR Exploration
Abstract
Reinforcement learning with verifiable rewards (RLVR) is often limited by insufficient exploration: difficult problems can yield uniformly incorrect rollout groups and therefore little learning signal. We show that such failures need not reflect missing capability. Instead, finite sampling often concentrates on a problem-specific dominant reasoning strategy while leaving alternative strategies already supported by the model unexplored. Moreover, the accessibility of these strategies evolves during RL: some are internalized into autonomous behavior, while others become difficult to elicit before being absorbed. Motivated by these observations, we introduce Problem--Strategy Rollout Allocation (PSRA), which treats unguided and strategy-conditioned prompts as competing exploration arms and uses Bayesian sequential allocation to direct a fixed rollout budget toward arms most likely to yield informative, non-saturated groups. A preservation objective keeps useful strategy-conditioned routes accessible while successful guided behaviors are transferred to the unguided policy. Across Qwen2.5 models from 1.5B to 7B and two RL training corpora, PSRA consistently improves reasoning performance, reduces dead saturation, strengthens out-of-distribution transfer, and maintains larger gains under increased inference budgets.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jin Cui, Xinyue Long, Boran Zhao, Pengju Ren, Hao Dong. 2026-09-27. The Model Knows Another Way: Strategy Switching for Effective RLVR Exploration. https://arxiv.org/abs/2609.33085
Cite the original work for its findings. Save a collection to share your selection of sources.