TY - RPRT TI - Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models AU - Zhipeng Chen AU - Xiaobo Qin AU - Youbin Wu AU - Yue Ling AU - Qinghao Ye AU - Wayne Xin Zhao AU - Guang Shi PY - 2025 UR - https://arxiv.org/abs/2508.10751 ID - 2508.10751 ER -