arXiv · 2606.24020
You Don't Need to Run Every Eval
Abstract
A modern model release reports scores on 40+ benchmarks and the same evaluations were run many more times before it: to track training progress, compare design choices, and select the checkpoint for the release. But do we need to run every eval? We compile a public score matrix of 84 frontier models on 133 benchmarks (2,604 cells, 23.3% filled) and find it is approximately rank-2: a model's scores across all benchmarks are largely determined by just two numbers. We confirm this in two ways: (i) scores hidden from the matrix are best recovered using two factors, and (ii) two factors already explain over 90% of the variation among models on the benchmarks they share. Building on this, we design BenchPress: a logit-space rank-2 matrix completion method that recovers held-out scores to within 4.6 points. Using BenchPress, we find a subset of five benchmarks {GPQA-D, HLE, Codeforces, MMLU-Pro, ARC-AGI-1} that can recover the rest of a model's public scorecard to within 3.93 points. For a tighter evaluation budget, a cheaper set {GPQA-D, MMLU-Pro, Aider Polyglot, MATH-500, AIME 2026} can predict a model's evals to within 4.55. Finally, we identify what affects prediction reliability and combine these factors with predictor disagreement to quantify when predictions can be trusted. We release the score matrix, the code for reproducing all of our experiments on Github.
Explore related subjects
Keep this discovery
Yuchen Zeng, Dimitris Papailiopoulos. 2026-06-22. You Don't Need to Run Every Eval. https://arxiv.org/abs/2606.24020
Cite the original work for its findings. Save a collection to share your selection of sources.