arXiv · 2609.33221
RMB: Reward Model Boosting Mitigates Reward Hacking
Abstract
Reinforcement Learning from Human Feedback (RLHF) is a powerful technique for aligning large language models (LLMs) with human preference. However, it often suffers from the reward hacking issue, where policy optimization improves the proxy reward model while actually degrading performance with respect to the true human preference, due to the imperfection of the proxy. To address this, we propose Reward Model Boosting (RMB), a novel approach that enhances the robustness and reliability of the reward signal for RLHF. RMB first trains a set of reward models with a diversity-promoting regularizer. This encourages each model to learn complementary aspects of the reward landscape. Then, RMB learns a lightweight aggregator in the principle of boosting to aggregate the outputs of the diverse reward models into a more accurate and robust reward signal. Our extensive experiments demonstrate that RMB significantly improves reward accuracy on both in-distribution and out-of-distribution datasets, substantially mitigating the reward hacking issue and ultimately improving RLHF performance.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jiabin Fan, Dezhi Ye, Yongchang Hao, Lili Mou. 2026-09-27. RMB: Reward Model Boosting Mitigates Reward Hacking. https://arxiv.org/abs/2609.33221
Cite the original work for its findings. Save a collection to share your selection of sources.