SearcharxivSearch

arXiv subjects

Qiyao Yu

Publications and source records attributed to Qiyao Yu.

3 recordsLinked to original sources

LiveMathematicianBench: A Live Benchmark for Mathematician-Level Reasoning with Proof Sketches

Mathematical reasoning is a hallmark of human intelligence, and whether large language models (LLMs) can meaningfully perform it remains a central question in artificial intelligence and cognitive science. As LLMs are increasingly integrated into scientific workflows, rigorous evaluation of their mathematical capabilities becomes a practical necessity. Existing benchmarks are limited by synthetic settings and data contamination. We present LiveMathematicianBench, a dynamic multiple-choice benchmark for research-level mathematical reasoning built from recent arXiv papers published after model training cutoffs. By grounding evaluation in newly published theorems, it provides a realistic testbed beyond memorized patterns. The benchmark introduces a thirteen-category logical taxonomy of theorem types (e.g., implication, equivalence, existence, uniqueness), enabling fine-grained evaluation across reasoning forms. It employs a proof-sketch-guided distractor pipeline that uses high-level proof strategies to construct plausible but invalid answer choices reflecting misleading proof directions, increasing sensitivity to genuine understanding over surface-level matching. We also introduce a substitution-resistant mechanism to distinguish answer recognition from substantive reasoning. Evaluation shows the benchmark is far from saturated: Gemini-3.1-pro-preview, the best model, achieves only 43.5%. Under substitution-resistant evaluation, accuracy drops sharply: GPT-5.4 scores highest at 30.6%, while Gemini-3.1-pro-preview falls to 17.6%, below the 20% random baseline. A dual-mode protocol reveals that proof-sketch access yields consistent accuracy gains, suggesting models can leverage high-level proof strategies for reasoning. Overall, LiveMathematicianBench offers a scalable, contamination-resistant testbed for studying research-level mathematical reasoning in LLMs.

cs.CL

On tori periods of Weil representations of unitary groups

We determine the restriction of Weil representations of unitary groups to maximal tori. In the local case, we show that the Weil representation contains a pair of compatible characters if and only if a root number condition holds. In the global case, we show that a torus period corresponding to a maximal anisotropic torus of the global theta lift of a character does not vanish if and only if the local condition is satisfied everywhere and a central value of an $L$-function does not vanish. Our proof makes use of the seesaw argument and of the well-known theta lifting results from $\operatorname{U}\left(1\right)$ to $\operatorname{U}\left(1\right)$. Our results are used in other papers to construct Arthur packets for $G_2$.

math.RT

On counting totally imaginary number fields

A number field is said to be a CM-number field if it is a totally imaginary quadratic extension of a totally real number field. We define a totally imaginary number field to be of CM-type if it contains a CM-subfield, and of TR-type if it does not contain a CM-subfield. For quartic totally imaginary number fields when ordered by discriminant, we show that about 69.95% are of TR-type and about 33.05% are of CM-type. For a sextic totally imaginary number field we classify its type in terms of its Galois group and possibly some additional information about the location of complex conjugation in the Galois group.

math.NT