SearcharxivSearch

arXiv subjects

Tomos Parry

Publications and source records attributed to Tomos Parry.

9 recordsLinked to original sources

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs

As Large Language Models (LLMs) saturate elementary benchmarks, the research frontier has shifted from generation to the reliability of automated evaluation. We demonstrate that standard "LLM-as-a-Judge" protocols suffer from a systematic Alignment Gap when applied to upper-undergraduate to early graduate level mathematics. To quantify this, we introduce QEDBench, the first large-scale dual-rubric alignment benchmark to systematically measure alignment with human experts on university-level math proofs by contrasting course-specific rubrics against expert common knowledge criteria. By deploying a dual-evaluation matrix (7 judges x 5 solvers) against 1,000+ hours of human evaluation, we reveal that certain frontier evaluators like Claude Opus 4.5, DeepSeek-V3, Qwen 2.5 Max, and Llama 4 Maverick exhibit significant positive bias (up to +0.18, +0.20, +0.30, +0.36 mean score inflation, respectively). Furthermore, we uncover a critical reasoning gap in the discrete domain: while Gemini 3.0 Pro achieves state-of-the-art performance (0.91 average human evaluation score), other reasoning models like GPT-5 Pro and Claude Sonnet 4.5 see their performance significantly degrade in discrete domains. Specifically, their average human evaluation scores drop to 0.72 and 0.63 in Discrete Math, and to 0.74 and 0.50 in Graph Theory. In addition to these research results, we also release QEDBench as a public benchmark for evaluating and improving AI judges. Our benchmark is publicly published at https://github.com/qqliu/Yale-QEDBench.

cs.LG

Primes in arithmetic progressions on average II

A deep conjecture of Montgomery and Soundararajan on the distribution of prime numbers in short intervals of length $h$ says that the third moment is bounded by $\ll h^{\frac {3}{2}-c}$ for some $c>0$. There is in the literature some conditional evidence towards this conjecture whilst in the first article to this series we gave the first instance of unconditional evidence in the form of a bound corresponding to $\ll h^{7/5+o(1)}$. In this article we push the exponent down to $\ll h^{1+o(1)}$ which more or less is expected to be best possible.

math.NT

Primes in arithmetic progressions on average I

Let $E_x(q,a)$ be the error term when counting primes in arithmetic progressions and let $M(Q)=\sum_{q\leq Q}\phi(q)\sum_{a=1}^qE_x(q,a)^3$. We show that $M(Q)<<Q^3(x/Q)^{7/5}$ for large $Q$ close to $x$ (in the usual BDH sense) thereby showing that sign changes in the error give power saving cancellation past the expected $\sqrt {x/q}$ heuristic.

math.NT

The distribution of $d_4(n)$ in arithmetic progressions

We use the Petrow-Young [10] subconvexity bound for Dirichlet $L$-functions to show that $d_4(n)$ has exponent of distribution $4/7$ when we allow an average over $a$ mod $q$, thereby giving an equidistribution result for $d_4(n)$ which goes past the $1/2$ barrier for the first time.

math.NT

A Montgomery-Hooley theorem for the k-fold divisor function

Let $d_k(n)$ denote the $k$-fold divisor function. For a wide range of large $q$ the expected bound $$\sum_{n\leq x\atop {n\equiv a(q)}}d_k(n)-\text { main term }\approx \sqrt {x/q}$$ is shown to be true in an average sense -- for all $k$. This generalises the work of Pongsriiam and Vaughan [15] who studied $k=2$, and answers the work of Rodgers and Soundararajan [17], who used the asymptotic large sieve to study a smoothed version of the problem. We use a circle method approach as developed by Goldston and Vaughan [7] to study the unsmoothed problem.

math.NT