TY - RPRT TI - Improving Policy Exploitation in Online Reinforcement Learning with Instant Retrospect Action AU - Gong Gao AU - Weidong Zhao AU - Xianhui Liu AU - Ning Jia PY - 2026 DO - 10.1016/j.neunet.2026.108667 UR - https://arxiv.org/abs/2601.19720 ID - 2601.19720 ER -