arXiv · 2609.29140
Sharp Limits for Honest Uncertainty in Hard-Budget Repeated Evaluation
Abstract
Repeated evaluation can estimate a benchmark score accurately while still requiring replication to certify narrow uncertainty. We characterize that requirement on a fixed grid of $M$ tasks with $L$ binary paths per task under the hard budget $(M+t)K$, where each path costs at most $K$ responses or episodes. For fixed $L \ge 3$ and $0 < α\le 1/12$, the optimal expected width on the worst pure cohort is $Θ_{α,L}([M(t+1)]^{-1/2})$ when every task is observed and $Θ_{α,L}([M(t+\sqrt{M})]^{-1/2})$ when omission is allowed. The lower bounds cover adaptive hard-budget policies, and fixed random-subset designs attain both rates through disagreement certificates. A joint mean/disagreement interval turns the task-covering law into practical finite-budget inference. In an equal-budget LiveCodeBench replay with 16 models, 880 tasks, and five outputs per task, the task-covering design reduces median point-estimation MSE by 87.0\% relative to pooled uniform sampling, while the Joint certificate produces narrower confidence intervals in 15/16 panels and reduces median interval width by 30.6\%. Finite-regime analyses identify task coverage as the effective choice at the evaluated scale and characterize how cohort size and within-task agreement determine the useful operating region. Together, the sharp laws and fixed-budget evidence make replication and task coverage explicit design variables for information-efficient repeated evaluation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yezhou Cheng, Runjia Du, Zeming Liu, Qibai Chen, Hang Lyu, Yilan Wei, Yankai Zeng, Bojun Lin. 2026-09-24. Sharp Limits for Honest Uncertainty in Hard-Budget Repeated Evaluation. https://arxiv.org/abs/2609.29140
Cite the original work for its findings. Save a collection to share your selection of sources.