arXiv · 2307.08526
Image Captions are Natural Prompts for Text-to-Image Models
Abstract
With the rapid development of Artificial Intelligence Generated Content (AIGC), it has become a common practice to train models on synthetic data due to data-scarcity and privacy leakage problems. Owing to massive and diverse information conveyed in real images, it is challenging for text-to-image generative models to synthesize informative training data with hand-crafted prompts. Considering the impressive ability of large generative models, could such models directly synthesize good training images for prediction tasks with proper prompts? We offer an affirmative response to this question by proposing a simple yet effective method, validated through ImageNet classification. Specifically, we caption each real image with the advanced captioning model to obtain informative and faithful prompts that extract class-relevant information and clarify the polysemy of class names. The image captions and class names are concatenated to prompt generative models for training image synthesis. We show that this simple caption incorporation significantly boosts the informativeness of synthetic data therefore enhancing downstream model generalization. More importantly, besides improvements in data augmentation and privacy preservation, our experiments demonstrate that synthesized images can exceed real data in terms of out-of-distribution robustness.
Explore related subjects
Keep this discovery
Shiye Lei, Hao Chen, Sen Zhang, Bo Zhao, Dacheng Tao. 2023-07-17. Image Captions are Natural Prompts for Text-to-Image Models. https://doi.org/10.1007/s11263-025-02436-0
Cite the original work for its findings. Save a collection to share your selection of sources.