The Flaw of Averages: Measuring Benchmark-Level Distributional Robustness
Benchmarks are central to measuring progress in language models, but aggregate scores can obscure substantial variation across subdomains, making models appear broadly competent despite concentrated strengths and weaknesses. We study this issue as benchmark-level distributional robustness: whether aggregate scores faithfully reflect performance across benchmark subdomains. We operationalize this notion with benchmark Harmony, an entropy-based measure of how uniformly model performance is distributed across subdomains. Measuring Harmony on 19 language model benchmarks across five model families, we find substantial variation in benchmark-level distributional robustness. Low-Harmony benchmarks are more likely to yield aggregate scores that overstate broad competence, whereas high-Harmony benchmarks provide more representative summaries of model capability. Rebalancing benchmarks by pruning overrepresented subdomains to increase Harmony substantially shifts aggregate scores for low-Harmony benchmarks, but leaves high-Harmony benchmarks comparatively stable. For example, while BoolQ remains comparatively stable as Harmony increases, PubMedQA, which evaluates performance in a medically consequential domain, exhibits substantial, often statistically significant, shifts in aggregate accuracy. Together, these findings show that aggregate scores can misrepresent broad competence when performance is unevenly distributed. We therefore recommend reporting benchmark Harmony alongside aggregate accuracy as a diagnostic of benchmark representativeness when interpreting claims about broad model competence.