TY - RPRT TI - Legal Mathematical Reasoning with LLMs: Procedural Alignment through Two-Stage Reinforcement Learning AU - Kepu Zhang AU - Guofu Xie AU - Weijie Yu AU - Mingyue Xu AU - Xu Tang AU - Yaxin Li AU - Jun Xu PY - 2025 UR - https://arxiv.org/abs/2504.02590 ID - 2504.02590 ER -