arXiv · 2609.31857
LLM Judge Validation Under Sparse Overlap: From Inference to Design
Abstract
Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this \emph{overlap sparsity} is the first-order determinant of wrong deployment decisions: at 5\% pairwise overlap, wrong-decision rates reach 25\% and the probability of selecting the wrong best judge among ten candidates is 65\%. The two actionable levers are overlap \emph{quantity} and \emph{allocation}. For quantity, we derive a minimum-overlap formula showing $ρ\geq 0.25$ suffices for non-borderline judges while borderline cases remain fundamentally hard. For allocation, a zero-cost stratified scheme halves false-rejection rates relative to random sampling when strata are informative. We validate on 10 LLM judges across four evaluation matrices spanning visual assessment, causal reasoning, and summarization.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Junxuan Li, Arko Mukherjee, Soumyabrata Pal. 2026-09-25. LLM Judge Validation Under Sparse Overlap: From Inference to Design. https://arxiv.org/abs/2609.31857
Cite the original work for its findings. Save a collection to share your selection of sources.