TY - RPRT TI - AMFT: Aligning LLM Reasoners by Meta-Learning the Optimal Imitation-Exploration Balance AU - Lixuan He AU - Jie Feng AU - Yong Li PY - 2025 UR - https://arxiv.org/abs/2508.06944 ID - 2508.06944 ER -