arXiv · 2609.32007
Interactive Proofs of Proximity for Model Evaluation
Abstract
We study interactive proofs of proximity (IPPs) for model evaluation, where a resource-limited verifier interacts with an untrusted prover, typically the model owner, to certify statistical properties of a model under an unknown input distribution. Our formulation separates sampling the input distribution from querying the model and evaluating its output; distinguishes real audit data (black-box sampling) from generated data (chosen-randomness, or gray-box, access to the sampler); and allows the prover and verifier to use different evaluators. We focus on doubly-sublinear IPPs, where both the verifier and honest prover use sublinear resources, and on (weighted) Hamming weight properties. For ordinary Hamming weight, we give a tolerant doubly-sublinear IPP. For completeness and soundness radii $\varepsilon_c<\varepsilon_f$ and gap $g=\varepsilon_f-\varepsilon_c$, a logarithmic-round instantiation uses $\widetilde{O}(1/g)$ verifier queries and $O(1/g^2)$ honest-prover queries, improving the cubic dependence of Amir, Goldreich, and Rothblum (ITCS 2025). We prove matching query lower bounds up to polylogarithmic factors. For distribution-weighted Hamming weight, black-box sampling requires $Θ(1/g^2)$ verifier samples but only $\widetilde{O}(1/g)$ evaluations; the quadratic sample complexity is necessary in the interior regime. With chosen-randomness access, the problem reduces to ordinary Hamming weight, yielding $\widetilde{O}(1/g)$ calls and evaluations. If the parties' evaluators disagree arbitrarily on a $ρ$-fraction of the distribution and by at most $γ$ elsewhere, our protocols remain doubly sublinear whenever $g>2κ$, where $κ=ρ+(1-ρ)γ$. Applications include auditing accuracy, group fairness, calibration, harmlessness, usefulness, and average-case robustness.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Geoffroy Couteau, Nikolas Melissaris, Tamara Paris. 2026-09-25. Interactive Proofs of Proximity for Model Evaluation. https://arxiv.org/abs/2609.32007
Cite the original work for its findings. Save a collection to share your selection of sources.