TY - RPRT TI - RL2ML: Finite-Rollout Surrogate Objectives from Reinforcement Learning to Maximum Likelihood AU - Yifu Zheng PY - 2026 UR - https://arxiv.org/abs/2605.30154 ID - 2605.30154 ER -