TY - RPRT TI - Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning AU - Yang Xu AU - Swetha Ganesh AU - Vaneet Aggarwal PY - 2026 UR - https://arxiv.org/abs/2506.07040 ID - 2506.07040 ER -