TY - RPRT TI - L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning AU - Pranjal Aggarwal AU - Sean Welleck PY - 2025 UR - https://arxiv.org/abs/2503.04697 ID - 2503.04697 ER -