TY - RPRT TI - MCPO: Mastery-Consolidated Policy Optimization for Large Reasoning Models AU - Zhaokang Liao AU - Yingguo Gao AU - Yi Yang AU - Yongheng Hu AU - Jingting Ding PY - 2026 UR - https://arxiv.org/abs/2604.16972 ID - 2604.16972 ER -