TY - RPRT TI - Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM Reasoning AU - Minwu Kim AU - Anubhav Shrestha AU - Safal Shrestha AU - Aadim Nepal AU - Keith Ross PY - 2025 UR - https://arxiv.org/abs/2505.14216 ID - 2505.14216 ER -