TY - RPRT TI - Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models AU - Guanting Dong AU - Keming Lu AU - Chengpeng Li AU - Tingyu Xia AU - Bowen Yu AU - Chang Zhou AU - Jingren Zhou PY - 2024 UR - https://arxiv.org/abs/2406.13542 ID - 2406.13542 ER -