TY - RPRT TI - Iterative Length-Regularized Direct Preference Optimization: A Case Study on Improving 7B Language Models to GPT-4 Level AU - Jie Liu AU - Zhanhui Zhou AU - Jiaheng Liu AU - Xingyuan Bu AU - Chao Yang AU - Han-Sen Zhong AU - Wanli Ouyang PY - 2024 UR - https://arxiv.org/abs/2406.11817 ID - 2406.11817 ER -