TY - RPRT TI - On-Policy Deep Reinforcement Learning for the Average-Reward Criterion AU - Yiming Zhang AU - Keith W. Ross PY - 2021 UR - https://arxiv.org/abs/2106.07329 ID - 2106.07329 ER -