arXiv · 2606.21077
OTTER: A Red-Teaming System for Toxicity-Evading Jailbreak Prompt Optimization
Abstract
Production LLMs increasingly rely on toxicity-based moderation filters as a primary defense, assuming that harmful intent correlates with toxic surface wording. We show this assumption is fundamentally brittle: surface toxicity and adversarial intent can be decoupled by replacing as few as five tokens. We present OTTER (Obfuscated Toxicity-Evading Token Evolution for Rewriting), a black-box red-teaming framework requiring only standard API access, directly targeting the practical constraints of industry security audits. Evaluated on 457 AdvBench prompts across four GPT models, OTTER raises average ASR from 7.0% to 84.0%. We further provide the first quantitative analysis of the toxicity--bypass relationship and a per-category breakdown, translating our findings into actionable recommendations for classifier hardening in production deployments.
Explore related subjects
Keep this discovery
Jerry Wang, Hsin-Ling Hsu, Yi-Cheng Lai, Nai-Chia Chen, Fang Yu. 2026-06-19. OTTER: A Red-Teaming System for Toxicity-Evading Jailbreak Prompt Optimization. https://arxiv.org/abs/2606.21077
Cite the original work for its findings. Save a collection to share your selection of sources.