arXiv · 2409.11904
Finding the Subjective Truth: Collecting 2 Million Votes for Comprehensive Gen-AI Model Evaluation
Abstract
Efficiently evaluating the performance of text-to-image models is difficult as it inherently requires subjective judgment and human preference, making it hard to compare different models and quantify the state of the art. Leveraging Rapidata's technology, we present an efficient annotation framework that sources human feedback from a diverse, global pool of annotators. Our study collected over 2 million annotations across 4,512 images, evaluating four prominent models (DALL-E 3, Flux.1, MidJourney, and Stable Diffusion) on style preference, coherence, and text-to-image alignment. We demonstrate that our approach makes it feasible to comprehensively rank image generation models based on a vast pool of annotators and show that the diverse annotator demographics reflect the world population, significantly decreasing the risk of biases.
Explore related subjects
Keep this discovery
Dimitrios Christodoulou, Mads Kuhlmann-Jørgensen. 2024-09-18. Finding the Subjective Truth: Collecting 2 Million Votes for Comprehensive Gen-AI Model Evaluation. https://arxiv.org/abs/2409.11904
Cite the original work for its findings. Save a collection to share your selection of sources.