SearcharxivSearch

arXiv subjects

Meng Gao

Publications and source records attributed to Meng Gao.

At least 19 recordsLinked to original sources

Quantum-accurate atomistic modeling of enzyme catalysis using a machine learned potential

Electronic rearrangements associated with bond forming/breaking in catalytic enzymes require quantum mechanical (QM) treatment beyond classical molecular mechanics (MM). Hybrid QM/MM methods enable tractable simulations but require system-specific setup and are sensitive to the QM region choice and treatment of the QM/MM interface. We demonstrate quantum-accurate treatment of all-atom, complete enzymes in explicit solvent comprising up to 54k atoms and 1 microsecond of total simulation time using the machine-learned interatomic potential (MLIP) eSEN-omol. We reproduce experimental barrier trends for Claisen rearrangement in chorismate mutase, resolve critical intermediate states in PETase catalyzed polymer depolymerization, and distinguish mechanistic alternatives for metal-activated phosphoryl transfer in nucleoside diphosphate kinase. We realize 1000x speedups relative to typical QM/MM calculations without system-specific tuning. These results establish MLIPs as a practical route to QM-accurate simulations of enzyme catalysis.

physics.chem-ph

Orbital occupation selects structural dimensionality in binary transition-metal oxides

Orbital-lattice coupling in transition-metal oxides is usually discussed within a given bonding framework, where orbital occupation is intertwined with local coordination, strain, or symmetry breaking. Here, we show that orbital occupation can also select the bonding framework itself, thereby determining structural dimensionality. Using first-principles calculations, we identify CrO as a prototype in which the single active $3d\,e_g$ electron of high-spin Cr$^{2+}$ gives rise to two competing orbital-structure states. The $d_{x^2-y^2}$ occupation favors a three-dimensionally connected covalent phase, whereas the $d_{z^2}$ occupation stabilizes a weakly coupled layered phase. Constrained-occupation calculations show that increasing the $d_{z^2}$ filling continuously contracts the in-plane lattice while expanding the structure along the layer normal. The two phases exhibit distinct magnetic ground states and ferroelastic responses. Moreover, the layered phase is robust against exchange-correlation functional and on-site ($U$) variations, remains dynamically stable down to the monolayer limit, and has a low exfoliation energy of 46 meV/Angstrom^2. Extending the analysis across related $3d$ binary oxides reveals a filling-dependence relation between accessible orbital filling and the preference for 2D or 3D connected bonding motifs, providing a microscopic basis for exploring low-dimensional oxide materials.

cond-mat.mtrl-sci

AudioSpan: Spanning the Duration and Depth of Audio Comprehension

General audio comprehension now covers speech, sound, and music over durations from seconds to hours, driven by large audio-language models (LALMs) that are increasingly omni-modal. Yet the benchmarks that test them still rely on clips of seconds, where scores saturate and models converge; recent long-form efforts extend duration but evaluate long audio much as short clips are. We introduce AudioSpan, a benchmark that spans both duration and depth: it pairs audio from 10 minutes to over 2 hours with 3,240 questions across three cognitive levels, namely perception, understanding, and reasoning. Two paths supply the questions, differing in how question content is sourced and how ground truth is obtained. Native QA extracts questions from the audio's content, posing each as a multiple-choice item and an open-ended one graded by detailed rubrics. Anchor QA instead injects ground truth, planting acoustic anchors into the audio and building a perception-to-reasoning chain scored only to the first error. A fully automated pipeline constructs every item through structured captioning, QA generation, and adversarial critic feedback. Evaluating 12 LALMs on AudioSpan, we find the hard part comes before reasoning: distilling a few relevant facts from a long, redundant signal. This difficulty grows with audio length and falls hardest on perception, especially temporal grounding. AudioSpan is available at https://huggingface.co/datasets/holvan/AudioSpan.

cs.SD

Execution-grounded evaluation reveals hidden failures in language-model calculations for environmental science

Large language models are increasingly used for quantitative work in the environmental sciences, yet existing evaluations score only final answers, leaving calculation process unobserved. Here we introduce AtmosCoder-Bench, an execution-grounded benchmark that makes the calculation process visible. Built through a transferable semi-automated pipeline (436 problems, 3,910 variants, 7,029 graded quantities), every problem is validated to be unambiguous and human-solvable, with uniquely verifiable answers. We find that (i) multiple-choice formats inflate measured accuracy by at least 12 percentage points; (ii) many failures arise not from missing knowledge but from models failing to apply known formulas and constraints consistently throughout multi-step computation; and (iii) even frontier models remain weak when task-specific conditions invalidate familiar methods, often reverting to canonical solution patterns rather than adapting methods to the relevant physical regime, leaving expert oversight essential.

cs.CL

An efficient adaptive dimension selection algorithm for multidimensional probit graded response models

Multidimensional graded response models (MGRMs) are widely used for analyzing ordinal questionnaire data in psychological and educational assessments. A central challenge in applying these models is determining the number of latent dimensions. Conventional approaches usually fit multiple fixed-dimensional models and select among them using post-hoc criteria such as AIC, BIC, or cross-validation, which can be computationally demanding and ignore uncertainty in dimensionality during estimation. We develop an adaptive Bayesian dimension selection framework for probit MGRMs. Building on the cumulative shrinkage process, we assign a cumulative ordered spike-and-slab (COSS) prior to the column-specific variances of the item loading matrix. This prior induces increasing shrinkage across latent dimensions, allowing redundant dimensions to be shrunk toward zero while preserving flexibility for active dimensions. Albert--Chib latent response augmentation is used to handle the ordinal probit likelihood, yielding conditionally Gaussian updates for item loadings and latent traits. These updates are combined with Gibbs updates for threshold and shrinkage parameters in an efficient adaptive sampler. Simulation studies evaluate the proposed method in terms of dimension recovery, parameter estimation accuracy, and computational efficiency, with comparisons to conventional fixed-dimensional estimation and model selection procedures. The results show that the proposed approach accurately recovers the latent structure while avoiding repeated model fitting over multiple candidate dimensions. We further illustrate the method using real psychological assessment data, demonstrating its practical utility for uncovering interpretable latent structures in ordinal item responses.

stat.ML

Formation of holographic vortex in a rotating shell-shaped superfluid

We investigate the holographic superfluid dynamics subjected to external rotation on a spherical geometry. Through a linear perturbation analysis, we identify several dynamically unstable phases in the phase diagram, each characterized by distinct unstable modes. Employing fully nonlinear numerical simulations, we further demonstrate that these unstable modes generically drive the system into vortex-antivortex configurations with definite winding numbers, determined by the symmetry of the corresponding unstable modes.

hep-th

Skim and Skip: Hierarchical Adaptive Inference for Efficient Multimodal Retrieval

Universal multimodal retrieval (UMR) increasingly adopts multimodal large language models (MLLMs) as unified embedding backbones, but their strong retrieval performance comes at substantial inference cost. Existing methods typically rely on uniformly dense inference, where all input tokens are processed through the entire model and matched using the final-layer [EOS] representation. However, this paradigm overlooks two key forms of heterogeneity in multimodal retrieval: token contributions to the final retrieval embedding are highly uneven, and different queries require markedly different amounts of inference depth. To address this, we propose Skim and Skip (SAS), a hierarchical adaptive inference framework for efficient multimodal retrieval. SAS first performs token-level evidence selection to preserve only the input information most relevant to the final retrieval embedding, and then performs depth-adaptive inference to determine whether the current representation is already sufficient for reliable matching. Experiments on 12 MMEB retrieval tasks show that SAS retains about 99% of the dense baseline's average retrieval performance while achieving up to 1.64 times end-to-end speedup and up to 66.3% FLOPs reduction.

cs.IR

Speculative Sampling For Faster Molecular Dynamics

Molecular dynamics (MD) is a key tool for simulating the dynamical behavior of atomic systems. However, MD is inherently serial, which makes it difficult to increase single-system throughput with concurrent compute. To address this, we introduce Langevin Speculative Dynamics (LSD), a distributed and model-agnostic speculative sampler for accelerating MD without adding relative error. Inspired by speculative methods in language and diffusion modeling, LSD uses a draft model to propose fast simulation steps and verifies them in parallel with a slower target model, applying a transport map from the draft to the target distribution. We extend speculative sampling to second-order Langevin dynamics, derive the achievable speedup as a function of physical parameters, show that LSD generalizes across different systems and draft-target combinations with a 3-9x speedup, and confirm theoretically and empirically that LSD samples trajectories from its target model distribution.

cs.LG

A density version of quaternary Goldbach problem

Let $\mathcal{P}$ denote the set of all primes, and let $\underline\delta(P)$ denote the relative lower density of a subset $P$ in $\mathcal{P}$. Suppose that $P_1, P_2, P_3, P_4$ are four subsets of primes with $\underline\delta(P_1)+\underline\delta(P_2)>1$ and $ \underline\delta(P_3)+\underline\delta(P_4)>1.$ Then for every sufficiently large even integer $n$, there exist primes $p_i \in P_i$ $(i=1,2,3,4)$ such that $n=p_1+p_2+p_3+p_4$. The condition is the best possible.

math.NT

On the quadratic Waring-Goldbach problem with primes in Piatetski-Shapiro sets

In this paper, it is proved that, for any $\gamma_1,\gamma_2,\gamma_3,\gamma_4,\gamma_5\in(\frac{28}{29},1)$, every sufficiently large integer $n$ subject to $n\equiv5\pmod{24}$ can be represented as the sum of five squares of primes, i.e., \begin{equation*} n=p_1^2+p_2^2+p_3^2+p_4^2+p_5^2, \end{equation*} such that $p_i=\lfloor m_i^{1/\gamma_i}\rfloor$ for some $m_i\in\mathbb{N}^+$ for each $1\leqslant i\leqslant 5$. This result constitutes an improvement upon the previous result of Zhang and Zhai [29].

math.NT

AIVD: Adaptive Edge-Cloud Collaboration for Accurate and Efficient Industrial Visual Detection

Multimodal large language models (MLLMs) demonstrate exceptional capabilities in semantic understanding and visual reasoning, yet they still face challenges in precise object localization and resource-constrained edge-cloud deployment. To address this, this paper proposes the AIVD framework, which achieves unified precise localization and high-quality semantic generation through the collaboration between lightweight edge detectors and cloud-based MLLMs. To enhance the cloud MLLM's robustness against edge cropped-box noise and scenario variations, we design an efficient fine-tuning strategy with visual-semantic collaborative augmentation, significantly improving classification accuracy and semantic consistency. Furthermore, to maintain high throughput and low latency across heterogeneous edge devices and dynamic network conditions, we propose a heterogeneous resource-aware dynamic scheduling algorithm. Experimental results demonstrate that AIVD substantially reduces resource consumption while improving MLLM classification performance and semantic generation quality. The proposed scheduling strategy also achieves higher throughput and lower latency across diverse scenarios.

cs.CV

UPETrack: Unidirectional Position Estimation for Tracking Occluded Deformable Linear Objects

Real-time state tracking of Deformable Linear Objects (DLOs) is critical for enabling robotic manipulation of DLOs in industrial assembly, medical procedures, and daily-life applications. However, the high-dimensional configuration space, nonlinear dynamics, and frequent partial occlusions present fundamental barriers to robust real-time DLO tracking. To address these limitations, this study introduces UPETrack, a geometry-driven framework based on Unidirectional Position Estimation (UPE), which facilitates tracking without the requirement for physical modeling, virtual simulation, or visual markers. The framework operates in two phases: (1) visible segment tracking is based on a Gaussian Mixture Model (GMM) fitted via the Expectation Maximization (EM) algorithm, and (2) occlusion region prediction employing UPE algorithm we proposed. UPE leverages the geometric continuity inherent in DLO shapes and their temporal evolution patterns to derive a closed-form positional estimator through three principal mechanisms: (i) local linear combination displacement term, (ii) proximal linear constraint term, and (iii) historical curvature term. This analytical formulation allows efficient and stable estimation of occluded nodes through explicit linear combinations of geometric components, eliminating the need for additional iterative optimization. Experimental results demonstrate that UPETrack surpasses two state-of-the-art tracking algorithms, including TrackDLO and CDCPD2, in both positioning accuracy and computational efficiency.

cs.RO

Dynamical Phase Transition of Dark Solitons in Spherical Holographic Superfluids

In this paper, we employ, for the first time, the holographic gravity approach to investigate the dynamical stability of solitons in spherical superfluids. Transverse perturbations are applied to the background of spherical soliton configurations, and the collective excitation modes of the solitons are examined within the framework of linear analysis. Our study reveals the existence of two distinct unstable modes in the soliton configurations. Through fully nonlinear evolution schemes, the dynamical evolution and final states of the solitons are elucidated. The results demonstrate that the solitons exhibit both self-acceleration instability and snake instability at different temperatures, respectively. And we explore the corresponding temperature-dependent dynamical phase transitions. It is noteworthy that the dynamical behavior of spherical solitons is distinct from the planar case due to the presence of spherical curvature.

hep-th

Learning from the electronic structure of molecules across the periodic table

Machine-Learned Interatomic Potentials (MLIPs) require vast amounts of atomic structure data to learn forces and energies, and their performance continues to improve with training set size. Meanwhile, the even greater quantities of accompanying data in the Hamiltonian matrix H behind these datasets has so far gone unused for this purpose. Here, we provide a recipe for integrating the orbital interaction data within H towards training pipelines for atomic-level properties. We first introduce HELM ("Hamiltonian-trained Electronic-structure Learning for Molecules"), a state-of-the-art Hamiltonian prediction model which bridges the gap between Hamiltonian prediction and universal MLIPs by scaling to H of structures with 100+ atoms, high elemental diversity, and large basis sets including diffuse functions. To accompany HELM, we release a curated Hamiltonian matrix dataset, 'OMol_CSH_58k', with unprecedented elemental diversity (58 elements), molecular size (up to 150 atoms), and basis set (def2-TZVPD). Finally, we introduce 'Hamiltonian pretraining' as a method to extract meaningful descriptors of atomic environments even from a limited number atomic structures, and repurpose this shared embedding space to improve performance on energy-prediction in low-data regimes. Our results highlight the use of electronic interactions as a rich and transferable data source for representing chemical space.

physics.chem-ph

FastCSP: Accelerated Molecular Crystal Structure Prediction with Universal Model for Atoms

Molecular crystal structure prediction (CSP) is essential for applications in pharmaceuticals and organic electronics. However, CSP remains challenging and computationally intensive due to the need to explore a large search space with sub-kJ/mol accuracy to distinguish between competing polymorphs. While dispersion-inclusive density functional theory (DFT) offers the necessary precision, its computational cost is impractical for a large number of putative structures. Here, we present FastCSP, an open-source, end-to-end CSP workflow driven entirely by a single pretrained universal machine learning interatomic potential (MLIP), the Universal Model for Atoms (UMA), without any system-specific fine-tuning or DFT calculations. FastCSP integrates conformer generation, random structure generation via Genarris 3, geometry optimization, free energy evaluation, and conformer energy corrections, all powered by UMA. Benchmarked on 28 semi-rigid and 10 flexible molecules spanning 74 experimental polymorphs, FastCSP reliably recovers all known structures, ranking them within 9 kJ/mol of the global minimum. UMA reproduces dispersion-inclusive DFT results with high fidelity across chemically diverse compounds. Conformer corrections are particularly beneficial for flexible compounds with conformational polymorphism, such as ROY. UMA's accuracy, transferability, and computational cost thus eliminate the need for classical force fields in early-stage screening and DFT-based re-ranking in CSP workflows. The open-source release of the entire FastCSP workflow lowers the barrier to accessing CSP, enabling both pharmaceutical-grade and high-throughput polymorph screening within practical computational reach.

physics.chem-ph

Open Molecular Crystals 2025 (OMC25) Dataset and Models

The development of accurate and efficient machine learning models for predicting the structure and properties of molecular crystals has been hindered by the scarcity of publicly available datasets of structures with property labels. To address this challenge, we introduce the Open Molecular Crystals 2025 (OMC25) dataset, a collection of over 27 million molecular crystal structures containing 12 elements and up to 300 atoms in the unit cell. The dataset was generated from dispersion-inclusive density functional theory (DFT) relaxation trajectories of over 230,000 randomly generated molecular crystal structures of around 50,000 organic molecules. OMC25 comprises diverse chemical compounds capable of forming different intermolecular interactions and a wide range of crystal packing motifs. We provide detailed information on the dataset's construction, composition, structure, and properties. To demonstrate the quality and use cases of OMC25, we further trained and evaluated state-of-the-art open-source machine learning interatomic potentials. By making this dataset publicly available, we aim to accelerate the development of more accurate and efficient machine learning models for molecular crystals.

physics.chem-ph

UMA: A Family of Universal Models for Atoms

The ability to quickly and accurately compute properties from atomic simulations is critical for advancing a large number of applications in chemistry and materials science including drug discovery, energy storage, and semiconductor manufacturing. To address this need, Meta FAIR presents a family of Universal Models for Atoms (UMA), designed to push the frontier of speed, accuracy, and generalization. UMA models are trained on half a billion unique 3D atomic structures (the largest training runs to date) by compiling data across multiple chemical domains, e.g. molecules, materials, and catalysts. We develop empirical scaling laws to help understand how to increase model capacity alongside dataset size to achieve the best accuracy. The UMA small and medium models utilize a novel architectural design we refer to as mixture of linear experts that enables increasing model capacity without sacrificing speed. For example, UMA-medium has 1.4B parameters but only ~50M active parameters per atomic structure. We evaluate UMA models on a diverse set of applications across multiple domains and find that, remarkably, a single model without any fine-tuning can perform similarly or better than specialized models. We are releasing the UMA code, weights, and associated data to accelerate computational workflows and enable the community to continue to build increasingly capable AI models.

cs.LG

Violation of Luttinger's theorem in one-dimensional interacting fermions

Using the density matrix renormalization group method, we systematically investigate the evolution of the Luttinger integral in the one-dimensional generalized $t$-$V$ model as a function of filling and interaction strength, and identify three representative phases. In the weak-coupling regime, the zero-frequency Green's function exhibits a branch-cut structure at the Fermi momentum, and the Luttinger integral accurately reflects the particle density, indicating that the Luttinger theorem holds. As the interaction increases, the spectral weight near the Fermi momentum is gradually suppressed. Interestingly, in the strong coupling regime near half-filling, this singularity is progressively destroyed, accompanied by the emergence of momentum-space zeros in the real part of the Green's function, leading to a novel non-Fermi liquid metallic phase beyond the classic Luttinger liquid paradigm, where the Luttinger surface is no longer defined by a single singularity. While finite spectral weight remains at the original Fermi momentum, the singularity gradually diminishes. Meanwhile, zeros with negligible spectral weight appear away from this momentum, significantly affecting the integral. At exact half-filling, a single-particle gap opens, and the Green's function becomes nearly vanishing across the entire momentum space, indicating the complete suppression of low-energy electronic states consistent with the nature of an insulating charge-density-wave phase. These results suggest that the breakdown of the Luttinger theorem is not triggered by a single mechanism, but rather results from the interplay between interaction-driven evolution of excitation modes and the breaking of particle-hole symmetry, ultimately leading to a continuous reconstruction of the generalized Fermi surface from topologically protected to correlation-driven.

cond-mat.str-el