arXiv · 2605.21123
Linear-DPO: Linear Direct Preference Optimization for Diffusion and Flow-Matching Generative Models
Abstract
Direct Preference Optimization (DPO) is successful for alignment in LLMs but still faces challenges in text-to-image generation. Existing studies are confined to denoising diffusion models while overlooking flow-matching, and suffer from an objective mismatch when applying discrete NLP-based DPO to regression-based generative tasks.\ In this paper, we derive a generalized DPO objective that covers both diffusion and flow-matching via a unified reverse-time SDE framework, and point out from a gradient perspective that the standard DPO objective is suboptimal for text-to-image generation. Consequently, we propose Linear-DPO, which replaces the aggressive sigmoid-based utility function with a sustained linear utility and incorporates an EMA-updated reference model. Qualitative and quantitative experiments on diffusion models (SD1.5, SDXL) and flow-matching model (SD3-Medium) demonstrate the superiority of our approach over existing baselines.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Kesong Li, Yixuan Xu, Kuo-kun Tseng, Weiyi Lu, Kan Liu, Tao Lan. 2026-05-20. Linear-DPO: Linear Direct Preference Optimization for Diffusion and Flow-Matching Generative Models. https://arxiv.org/abs/2605.21123
Cite the original work for its findings. Save a collection to share your selection of sources.