arXiv · 2608.29921
Sleight of Word Benchmark: Can Language Models Notice If Their Own Output Was Tampered With?
Abstract
The output of a Language Model can be tampered with \emph{while} the model is writing it. A simple test can thus be constructed by evaluating the model's perception of this external perturbation. In this spirit, a simple benchmark is built in which a single word is consistently substituted with another in the generation process. We call this method \emph{Sleight of Word}. Two distinct axes are measured: metrics that relate to the model's surprise, as well as an evaluation of the textual reaction for 19 different open-weight language models.
Explore related subjects
Keep this discovery
Alberto Cetoli. 2026-08-30. Sleight of Word Benchmark: Can Language Models Notice If Their Own Output Was Tampered With?. https://arxiv.org/abs/2608.29921
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.