arXiv · 2602.17658
MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling
Abstract
Reward modeling is central to RLHF, RLAIF, and PPO-based alignment, but its reliability is often limited by scarce and heterogeneous human preference data. In this paper, we introduce MARS (Margin and Semantic-Aware Data Augmentation for Reward Modeling), an adaptive augmentation framework for controlled low-resource reward modeling. MARS allocates more augmentation to low-margin preference pairs and uses semantic-distance-based refinement to improve chosen-rejected contrast before generating synthetic preference samples. Across three preference datasets, two reward-model backbones, and downstream alignment evaluations, MARS improves average RewardBench performance and alignment win rates over uniform augmentation, WoN, and AdaBoost-style baselines. Ablations and independent-judge evaluations suggest that the gains are not solely explained by semantic refinement alone or GPT-4.1 judge coupling.
Explore related subjects
Keep this discovery
Payel Bhattacharjee, Osvaldo Simeone, Ravi Tandon. 2026-02-19. MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling. https://arxiv.org/abs/2602.17658
Cite the original work for its findings. Save a collection to share your selection of sources.