TY - RPRT TI - Mixed Preference Optimization: Reinforcement Learning with Data Selection and Better Reference Model AU - Qi Gou AU - Cam-Tu Nguyen PY - 2025 UR - https://arxiv.org/abs/2403.19443 ID - 2403.19443 ER -