arXiv · 2411.15516
When Image Generation Goes Wrong: A Safety Analysis of Stable Diffusion Models
Abstract
Text-to-image models are increasingly popular and impactful, yet concerns regarding their safety and fairness remain. This study investigates the ability of ten popular Stable Diffusion models to generate harmful images, including NSFW, violent, and personally sensitive material. We demonstrate that these models respond to harmful prompts by generating inappropriate content, which frequently displays troubling biases, such as the disproportionate portrayal of Black individuals in violent contexts. Our findings demonstrate a complete lack of any refusal behavior or safety measures in the models observed. We emphasize the importance of addressing this issue as image generation technologies continue to become more accessible and incorporated into everyday applications.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Matthias Schneider, Thilo Hagendorff. 2024-11-23. When Image Generation Goes Wrong: A Safety Analysis of Stable Diffusion Models. https://doi.org/10.1038/s41598-025-12032-4
Cite the original work for its findings. Save a collection to share your selection of sources.