arXiv · 2609.32482
From Outcomes to Strategies: Learning Strategy Utility for Mathematical Reasoning
Abstract
Reinforcement learning with verifiable rewards has substantially improved mathematical reasoning. However, terminal correctness alone provides limited insight into the quality of high-level strategies, such as theorem selection and subgoal decomposition, when considered separately from their subsequent execution. This paper studies strategy utility, which is defined as the likelihood that a strategy supports a correct downstream solution under a given executor. We introduce SURE, a framework for learning and leveraging relative strategy utility. In this framework, high-level strategies are separated from their detailed reasoning. Based on the pairwise preferences constructed from strategy-conditioned rollouts and teacher-generated contrasts, a Strategy Reward Model is learned to estimate relative strategy utility. During reinforcement learning, the frozen reward model reads only the extracted strategy, whose score is combined with the correctness and format rewards in a sequence-level GRPO objective. Compared with outcome-and-format GRPO baselines, experiments show that SURE improves average pass@1 by 1.87%, 2.64%, and 2.93% across three policy backbones. Our method also achieves competitive or better accuracy than stronger reward baselines while requiring substantially lower GRPO-stage compute.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ruikang Zhang, Xiao An, Xuli Shen, Jiaxing Sun, Xiaoyi Yu, Jin Zeng, Jiang Wu, Tong Lin. 2026-09-26. From Outcomes to Strategies: Learning Strategy Utility for Mathematical Reasoning. https://arxiv.org/abs/2609.32482
Cite the original work for its findings. Save a collection to share your selection of sources.