Searcharxiv⌕ Search

arXiv subjects

James V. Roggeveen

Publications and source records attributed to James V. Roggeveen.

6 recordsLinked to original sources

Measuring Progress in Reasoning Toward Mathematical Discovery with Automatic Verification

Can AI make progress on important, unsolved mathematical problems? Large language models are now capable of sophisticated mathematical and scientific reasoning, but whether they can perform novel research is still widely debated and underexplored. We introduce HorizonMath, a benchmark of 113 predominantly unsolved problems spanning eight domains in mathematics and the mathematical sciences, paired with an open-source evaluation framework for automated verification. Our benchmark targets the generator-verifier gap: problems where discovery is hard and requires meaningful mathematical insight, but verification is computationally straightforward. This contrasts with most existing research-level benchmarks, which instead rely on formal proof verification or manual review, both of which are expensive to scale. Because these solutions are unknown, HorizonMath is resistant to data contamination, and most state-of-the-art models score under 10%. Using this framework, we identify six novel solutions to research problems that either resolve previously open questions or improve on the best-known published results, with GPT-5.4 Pro and GPT-5.6 Sol each discovering three of these solutions. Across seven frontier model families, reasoning efficiency and behavior also vary substantially. We release HorizonMath as an open challenge and a growing community resource, where each verified solution is a candidate contribution to the mathematical literature.

cs.LG↗

Learning constitutive models and rheology from partial flow measurements

Constitutive laws relate fluid stress to deformation and underpin predictions of non-Newtonian behavior in industrial and biological fluids. Standard characterization relies on measurements in idealized flows that often miss physics relevant to complex geometries. Existing data-driven methods overfit sparse data, lack geometry portability, or presuppose constitutive forms. To unify measurement and constitutive discovery, we developed an end-to-end framework that leverages automatic differentiation through a full physics simulation. By embedding a frame-invariant tensor basis neural network (TBNN) within a differentiable non-Newtonian solver, we learn constitutive laws from any flow observable without presupposing a specific model, spanning generalized Newtonian, viscoelastic, and yield-stress behavior. Unlike coordinate-dependent methods, learning local material response enables accurate flow predictions in unseen geometries and conditions without retraining. We then distill the TBNN closure into symbolic form via automated model selection using the Bayesian Information Criterion, extracting interpretable physical parameters. This work establishes a foundation for comprehensive characterization of complex fluids directly within their operating environment ("digital rheometry") with broad applicability to constitutive discovery across engineering and the physical sciences.

physics.flu-dyn↗

CMT-Benchmark: A Benchmark for Condensed Matter Theory Built by Expert Researchers

Large language models (LLMs) have shown remarkable progress in coding and math problem-solving, but evaluation on advanced research-level problems in hard sciences remains scarce. To fill this gap, we present CMT-Benchmark, a dataset of 50 problems covering condensed matter theory (CMT) at the level of an expert researcher. Topics span analytical and computational approaches in quantum many-body, and classical statistical mechanics. The dataset was designed and verified by a panel of expert researchers from around the world. We built the dataset through a collaborative environment that challenges the panel to write and refine problems they would want a research assistant to solve, including Hartree-Fock, exact diagonalization, quantum/variational Monte Carlo, density matrix renormalization group (DMRG), quantum/classical statistical mechanics, and model building. We evaluate LLMs by programmatically checking solutions against expert-supplied ground truth. We developed machine-grading, including symbolic handling of non-commuting operators via normal ordering. They generalize across tasks too. Our evaluations show that frontier models struggle with all of the problems in the dataset, highlighting a gap in the physical reasoning skills of current LLMs. Notably, experts identified strategies for creating increasingly difficult problems by interacting with the LLMs and exploiting common failure modes. The best model, GPT5, solves 30\% of the problems; average across 17 models (GPT, Gemini, Claude, DeepSeek, Llama) is 11.4\pm2.1\%. Moreover, 18 problems are solved by none of the 17 models, and 26 by at most one. These unsolved problems span Quantum Monte Carlo, Variational Monte Carlo, and DMRG. Answers sometimes violate fundamental symmetries or have unphysical scaling dimensions. We believe this benchmark will guide development toward capable AI research assistants and tutors.

cs.LG↗

Meshless solutions of PDE inverse problems on irregular geometries

Solving inverse and optimization problems over solutions of nonlinear partial differential equations (PDEs) on complex spatial domains is a long-standing challenge. Here we introduce a method that parameterizes the solution using spectral bases on arbitrary spatiotemporal domains, whereby the basis is defined on a hyperrectangle containing the true domain. We find the coefficients of the basis expansion by solving an optimization problem whereby both the equations, the boundary conditions and any optimization targets are enforced by a loss function, building on a key idea from Physics-Informed Neural Networks (PINNs). Since the representation of the function natively has exponential convergence, so does the solution of the optimization problem, as long as it can be solved efficiently. We find empirically that the optimization protocols developed for machine learning find solutions with exponential convergence on a wide range of equations. The method naturally allows for the incorporation of data assimilation by including additional terms in the loss function, and for the efficient solution of optimization problems over the PDE solutions.

math.NA↗

HARDMath2: A Benchmark for Applied Mathematics Built by Students as Part of a Graduate Class

Large language models (LLMs) have shown remarkable progress in mathematical problem-solving, but evaluation has largely focused on problems that have exact analytical solutions or involve formal proofs, often overlooking approximation-based problems ubiquitous in applied science and engineering. To fill this gap, we build on prior work and present HARDMath2, a dataset of 211 original problems covering the core topics in an introductory graduate applied math class, including boundary-layer analysis, WKB methods, asymptotic solutions of nonlinear partial differential equations, and the asymptotics of oscillatory integrals. This dataset was designed and verified by the students and instructors of a core graduate applied mathematics course at Harvard. We build the dataset through a novel collaborative environment that challenges students to write and refine difficult problems consistent with the class syllabus, peer-validate solutions, test different models, and automatically check LLM-generated solutions against their own answers and numerical ground truths. Evaluation results show that leading frontier models still struggle with many of the problems in the dataset, highlighting a gap in the mathematical reasoning skills of current LLMs. Importantly, students identified strategies to create increasingly difficult problems by interacting with the models and exploiting common failure modes. This back-and-forth with the models not only resulted in a richer and more challenging benchmark but also led to qualitative improvements in the students' understanding of the course material, which is increasingly important as we enter an age where state-of-the-art language models can solve many challenging problems across a wide domain of fields.

cs.LG↗

Transport of a passive scalar in wide channels with surface topography

We generalize classical dispersion theory for a passive scalar to derive an asymptotic long-time convection-diffusion equation for a solute suspended in a wide, structured channel and subject to a steady low-Reynolds-number shear flow. Our theory, valid for small roughness amplitudes of the channel, holds for general surface shapes expandable as a Fourier series. We determine an anisotropic dispersion tensor, which depends on the characteristic wavelengths and amplitude of the surface structure. For surfaces whose corrugations are tilted with respect to the applied flow direction, we find that dispersion along the principal direction (i.e., the principal eigenvector of the dispersion tensor) is at an angle to the main flow direction and becomes enhanced relative to classical Taylor dispersion. In contrast, dispersion perpendicular to it can decrease compared to the short-time diffusivity of the particles. Furthermore, for an arbitrary surface shape represented in terms of a Fourier decomposition, we find that each Fourier mode contributes at leading order a linearly-independent correction to the classical Taylor dispersion tensor.

physics.flu-dyn↗