arXiv · 2502.06738
Resurrecting saturated LLM benchmarks with adversarial encoding
Abstract
Recent work showed that small changes in benchmark questions can reduce LLMs' reasoning and recall. We explore two such changes: pairing questions and adding more answer options, on three benchmarks: WMDP-bio, GPQA, and MMLU variants. We find that for more capable models, these predictably reduce performance, essentially heightening the performance ceiling of a benchmark and unsaturating it again. We suggest this approach can resurrect old benchmarks.
Explore related subjects
Keep this discovery
Igor Ivanov, Dmitrii Volkov. 2025-02-10. Resurrecting saturated LLM benchmarks with adversarial encoding. https://arxiv.org/abs/2502.06738
Cite the original work for its findings. Save a collection to share your selection of sources.