arXiv · 2508.05775
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
Abstract
Large Language Models (LLMs) have revolutionized content creation across digital platforms, offering unprecedented capabilities in natural language generation and understanding. Meanwhile, they pose risks by inadvertently producing toxic, offensive, or biased content. This dual role of LLMs, both as powerful tools for text generation and as potential sources of harmful language, presents a pressing sociotechnical challenge. In this survey, we systematically review recent studies encompassing unintentional toxicity, adversarial jailbreak attacks, and comprehensive mitigation strategies. We explore LLMs' dual role as both generators of harm and enablers of safety through detection, classification, content moderation, and prevention. We propose a unified taxonomy of LLM-related harms and defenses, analyze emerging multimodal and LLM-assisted jailbreak strategies, and assess mitigation efforts, including reinforcement learning with human feedback (RLHF), prompt engineering, and safety alignment. Our review highlights the evolving landscape of LLM safety and identifies limitations in current evaluation methodologies. Ultimately, our review outlines future research directions to guide the development of robust and ethically aligned language technologies.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chi Zhang, Changjia Zhu, Junjie Xiong, Xiaoran Xu, Lingyao Li, Yao Liu, Zhuo Lu. 2025-08-07. Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM. https://arxiv.org/abs/2508.05775
Cite the original work for its findings. Save a collection to share your selection of sources.