SearcharxivSearch

arXiv subjects

Zehong Zhao

Publications and source records attributed to Zehong Zhao.

4 recordsLinked to original sources

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language. Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. In this work, we introduce VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable. 1) Task scaling. VBVR-Pro turns visual reasoning into a controlled task space of 300 procedurally generated tasks. Models trained on VBVR-Pro show strong transfer beyond the proposed suite across seven external visual reasoning benchmarks such as RISE-Video, MME-CoF-Pro, and BabyVision. 2) Verifiable rewards. VBVR-Pro provides verifiable reward scorers for task-grounded evaluation. Through a systematic study of leading MLLMs as judges, we identify recurring failure modes of the prevalent VLM-as-a-judge paradigm. In contrast, the proposed scorers are grounded in deterministic, task-specific rules, achieve fine-grained alignment with human judgments. Importantly, they serve as reliable reward signals for large-scale multi-task reinforcement learning and demonstrate stronger post-RL performance across visual reasoning tasks. 3) Mechanism study. VBVR-Pro enables controlled modality studies across more than 30 image, video, and interleaved generators. Our analysis shows that video generation remains strongest for tasks requiring persistent spatiotemporal state tracking, while interleaved generation provides a compute-efficient alternative. Critically, ablations and probing suggest the presence of vision-native trajectories that are crucial to visual reasoning. We release all data, models, scorers, and code.

cs.CV

A Very Big Video Reasoning Suite

Rapid progress in video models has largely focused on visual quality, leaving their reasoning capabilities underexplored. Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can naturally capture, enabling intuitive reasoning over spatiotemporal structure such as continuity, interaction, and causality. However, systematically studying video reasoning and its scaling behavior is hindered by the lack of large-scale training data. To address this gap, we introduce the Very Big Video Reasoning (VBVR) Dataset, an unprecedentedly large-scale resource spanning 200 curated reasoning tasks following a principled taxonomy and over one million video clips, approximately three orders of magnitude larger than existing datasets. We further present VBVR-Bench, a verifiable evaluation framework that moves beyond model-based judging by incorporating rule-based, human-aligned scorers, enabling reproducible and interpretable diagnosis of video reasoning capabilities. Leveraging the VBVR suite, we conduct one of the first large-scale scaling studies of video reasoning and observe early signs of emergent generalization to unseen reasoning tasks. Together, VBVR lays a foundation for the next stage of research in generalizable video reasoning. The data, benchmark toolkit, and models are publicly available at https://video-reason.com/?v=vbvr .

cs.CV

Satisfiability Phase Transtion for Random Quantum 3XOR Games

Recent results showed it was possible to determine if a modest size 3XOR game has a perfect quantum strategy. We build on these and give an explicit polynomial time algorithm which constructs such a perfect strategy or refutes its existence. This new tool lets us numerically study the behavior of randomly generated 3XOR games with large numbers of questions. A key issue is: how common are pseudotelephathy games (games with perfect quantum strategies but no perfect classical strategies)? Our experiments strongly indicate that the probability of a randomly generated game being pseudotelpathic stays far from 1, indeed it is bounded below 0.15. We also find strong evidence that randomly generated 3XOR games undergo both a quantum and classical "phase transition", transitioning from almost certainly perfect to almost certainly imperfect as the ratio of number of clauses ($m$) to number of questions ($n$) increases. The locations of these two phase transitions appear to coincide at $m/n \approx 2.74$.

quant-ph

Harmonic bases for generalized coinvariant algebras

Let $k \leq n$ be nonnegative integers and let $λ$ be a partition of $k$. S. Griffin recently introduced a quotient $R_{n,λ}$ of the polynomial ring $\mathbb{Q}[x_1, \dots, x_n]$ in $n$ variables which simultaneously generalizes the Delta Conjecture coinvariant rings of Haglund-Rhoades-Shimozono and the cohomology rings of Springer fibers studied by Tanisaki and Garsia-Procesi. We describe the space $V_{n,λ}$ of harmonics attached to $R_{n,λ}$ and produce a harmonic basis of $R_{n,λ}$ indexed by certain ordered set partitions $\mathcal{OP}_{n,λ}$. The combinatorics of this basis is governed by a new extension of the {\em Lehmer code} of a permutation to $\mathcal{OP}_{n, λ}$.

math.CO