TY - RPRT TI - Offline Reinforcement Learning for LLM Multi-Step Reasoning AU - Huaijie Wang AU - Shibo Hao AU - Hanze Dong AU - Shenao Zhang AU - Yilin Bao AU - Ziran Yang AU - Yi Wu PY - 2024 UR - https://arxiv.org/abs/2412.16145 ID - 2412.16145 ER -