arXiv · 2610.01548
Range-GRPO: Policy Optimization via Pairwise Relations among Reward Intervals
Abstract
As the use of large language models (LLMs) expands, post-training has become increasingly important for adapting them to downstream tasks. However, obtaining reliable supervision remains costly, especially in domains without reference answers or executable verifiers. LLM-as-a-Judge provides scalable pseudo-rewards for unlabeled responses, but a single point score does not explicitly represent reward uncertainty. This motivates representing pseudo-rewards as conformally calibrated reward ranges. We propose Range-GRPO, a semi-supervised post-training framework that combines limited labeled data with unlabeled prompts. In Group Relative Policy Optimization (GRPO), learning signals depend on relative reward comparisons within each rollout group. The proposed objective compares reward ranges pairwise rather than reducing them to point rewards, allowing interval uncertainty to affect both the magnitude and direction of these signals. Our theoretical analysis characterizes this distinction and shows that the proposed objective recovers the Dr.GRPO advantage when all reward ranges collapse to points. Empirically, Range-GRPO achieves the highest in-distribution and out-of-distribution average performance among the evaluated semi-supervised methods while requiring fewer training resources.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ryunyi Lee, Kangjun Noh, Somin Kim, Heedong Kim, Kyungwoo Song. 2026-10-01. Range-GRPO: Policy Optimization via Pairwise Relations among Reward Intervals. https://arxiv.org/abs/2610.01548
Cite the original work for its findings. Save a collection to share your selection of sources.