arXiv · 2508.14070
Special-Character Adversarial Attacks on Open-Source Language Model
Abstract
Large language models (LLMs) have achieved remarkable performance across diverse natural language processing tasks, yet their vulnerability to character-level adversarial manipulations presents significant security challenges for real-world deployments. This paper presents a study of different special character attacks including unicode, homoglyph, structural, and textual encoding attacks aimed at bypassing safety mechanisms. We evaluate seven prominent open-source models ranging from 3.8B to 32B parameters on 4,000+ attack attempts. These experiments reveal critical vulnerabilities across all model sizes, exposing failure modes that include successful jailbreaks, incoherent outputs, and unrelated hallucinations.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ephraiem Sarabamoun. 2025-08-12. Special-Character Adversarial Attacks on Open-Source Language Model. https://arxiv.org/abs/2508.14070
Cite the original work for its findings. Save a collection to share your selection of sources.