arXiv · 2506.08480
Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models
Abstract
Text-to-image models often struggle to generate images that precisely match textual prompts. Prior research has extensively studied the evaluation of image-text alignment in text-to-image generation. However, existing evaluations primarily focus on agreement with human assessments, neglecting other critical properties of a trustworthy evaluation framework. In this work, we first identify two key aspects that a reliable evaluation should address. We then empirically demonstrate that current mainstream evaluation frameworks fail to fully satisfy these properties across a diverse range of metrics and models. Finally, we propose recommendations for improving image-text alignment evaluation.
Explore related subjects
Keep this discovery
Huixuan Zhang, Xiaojun Wan. 2025-06-10. Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models. https://arxiv.org/abs/2506.08480
Cite the original work for its findings. Save a collection to share your selection of sources.