TY - RPRT TI - Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs AU - Nicolas Le Roux AU - Marc G. Bellemare AU - Jonathan Lebensold AU - Arnaud Bergeron AU - Joshua Greaves AU - Alex Fréchette AU - Carolyne Pelletier AU - Eric Thibodeau-Laufer AU - Sándor Toth AU - Sam Work PY - 2025 UR - https://arxiv.org/abs/2503.14286 ID - 2503.14286 ER -