arXiv · 2609.22183
CleanScore: Black-Box Benchmark Audits with Negative Controls and Sensitivity Bounds
Abstract
Public benchmark scores may reflect skill, prior exposure to the questions, or both, and for most models the training data are unknown. We present CleanScore, a black-box audit using scored outputs only. Each benchmark question becomes a parent item with one public form and two independently written fresh forms preserving its numbers, facts and answer. The audit reports an interval for the public-form advantage rather than a verdict, and a private negative-control bank with an explicit transport radius separates exposure from ordinary form mismatch. A registered controlled-exposure experiment detects planted exposure and stays quiet under fresh-form exposure. A registered audit of five open models on 200 GSM8K and 200 ARC-Challenge items finds no exposure-consistent advantage, bounding surface-form inflation below five points. Registered positive controls then bound what such a null can mean. Leaking an item raises accuracy on paraphrases the model never saw almost as much as on the leaked wording, leaving 52% to 110% of the effect invisible to a paraphrase audit. On ARC a planted 49-point advantage shows an observable gap of -0.020, and about 20 points survive rewriting stem and options, across four training seeds. A surface-form null bounds far less than the phrase contamination audit implies.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jeffery Opoku, David Banahene. 2026-08-27. CleanScore: Black-Box Benchmark Audits with Negative Controls and Sensitivity Bounds. https://arxiv.org/abs/2609.22183
Cite the original work for its findings. Save a collection to share your selection of sources.