TY - RPRT TI - Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning AU - Heyang Jiang AU - Henry Liu AU - Baharan Mirzasoleiman PY - 2026 UR - https://arxiv.org/abs/2607.22002 ID - 2607.22002 ER -