TY - RPRT TI - Refined Direct Preference Optimization with Synthetic Data for Behavioral Alignment of LLMs AU - VĂ­ctor Gallego PY - 2024 UR - https://arxiv.org/abs/2402.08005 ID - 2402.08005 ER -