arXiv · 2608.02681
Safety in Batches? Understanding and Mitigating Safety Failures in Batch Prompting
Abstract
Batch prompting is a practical inference strategy for large language models, but its safety implications remain underexplored. We show that the success of batch prompting for utility does not extend to safety: a harmful question that is reliably refused in isolation can elicit a harmful response when embedded in a batch of benign questions. We identify this as a distinct safety failure mode, not reducible to known vulnerabilities such as in-context learning or long-context effects, and analyze its causes from two complementary perspectives: alignment signal weakening and refusal signal dilution. Across widely used open-source and frontier commercial models, batch prompting consistently achieves high attack success rates as a simple black-box attack. We further show that batch-aware preference optimization effectively mitigates the vulnerability. These findings highlight a blind spot in current safety alignment and point to batch-aware alignment as a necessary step toward robust deployment.
Explore related subjects
Keep this discovery
Kihyun Kim, Hee-Seon Kim, Wonjun Lee, Changick Kim. 2026-08-03. Safety in Batches? Understanding and Mitigating Safety Failures in Batch Prompting. https://arxiv.org/abs/2608.02681
Cite the original work for its findings. Save a collection to share your selection of sources.