arXiv · 2601.07118
Reward-Preserving Attacks For Robust Reinforcement Learning
Abstract
Adversarial training in reinforcement learning (RL) is challenging because perturbations cascade through trajectories and compound over time, making fixed-strength attacks either overly destructive or too conservative. We propose reward-preserving attacks, which adapt adversarial strength so that an $\alpha$ fraction of the nominal-to-worst-case return gap remains achievable at each state. In deep RL, perturbation magnitudes $\eta$ are selected dynamically, using a learned critic $Q((s,a),\eta)$ that estimates the expected return of $\alpha$-reward-preserving rollouts. For intermediate values of $\alpha$, this adaptive training yields policies that are robust across a wide range of perturbation magnitudes while preserving nominal performance, outperforming fixed-radius and uniformly sampled-radius adversarial training.
Explore related subjects
Keep this discovery
Lucas Schott, Elies Gherbi, Hatem Hajri, Sylvain Lamprier. 2026-01-12. Reward-Preserving Attacks For Robust Reinforcement Learning. https://arxiv.org/abs/2601.07118
Cite the original work for its findings. Save a collection to share your selection of sources.