TY - RPRT TI - Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning AU - Zimu Lu AU - Aojun Zhou AU - Ke Wang AU - Houxing Ren AU - Weikang Shi AU - Junting Pan AU - Mingjie Zhan AU - Hongsheng Li PY - 2024 UR - https://arxiv.org/abs/2407.00782 ID - 2407.00782 ER -