TY - RPRT TI - Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards AU - Johannes Ackermann AU - Michael Noukhovitch AU - Takashi Ishida AU - Masashi Sugiyama PY - 2026 UR - https://arxiv.org/abs/2602.18037 ID - 2602.18037 ER -