TY - RPRT TI - $\textbf{PLUM}$: Improving Code LMs with Execution-Guided On-Policy Preference Learning Driven By Synthetic Test Cases AU - Dylan Zhang AU - Shizhe Diao AU - Xueyan Zou AU - Hao Peng PY - 2024 UR - https://arxiv.org/abs/2406.06887 ID - 2406.06887 ER -