TY - RPRT TI - On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes AU - Rishabh Agarwal AU - Nino Vieillard AU - Yongchao Zhou AU - Piotr Stanczyk AU - Sabela Ramos AU - Matthieu Geist AU - Olivier Bachem PY - 2024 UR - https://arxiv.org/abs/2306.13649 ID - 2306.13649 ER -