SearcharxivSearch

arXiv subjects

Tianyu Zhu

Publications and source records attributed to Tianyu Zhu.

At least 19 recordsLinked to original sources

Binarized High-Efficiency RAW Video Restoration and Beyond

RAW video restoration is fundamental to high-quality low-level perception and serves as the basis for a wide range of downstream vision applications. While binary neural networks (BNNs) enable efficient lightweight deployment for image enhancement, their deficiencies in modeling temporal coherence and activation value distributions hinder their effectiveness when applied to video scenarios. In this paper, we propose BinRVR, a binarized RAW video restoration framework that reduces computation and parameters by approximately 96% while incurring only about 4% performance degradation. Specifically, we present a Binarized Information Interaction Module (BIIM) to jointly model spatial and temporal information in an efficient and unified manner. Moreover, we develop a Distribution-Aware Binarized Convolution (DAB-Conv) that leverages the statistics of full-precision activations to mitigate quantization errors. The proposed framework further supports multi-bit quantization, enabling flexible accuracy-efficiency trade-offs across different hardware constraints. Extensive experiments demonstrate that our BinRVR achieves competitive performance compared with state-of-the-art binarized methods on RAW video restoration tasks, including low-light enhancement, denoising, deblurring, and super-resolution. We further explore the potential of our method on downstream video applications, including object detection and monocular depth estimation.

cs.CV

Resolving Finite-Size Errors in EOM-CCSD Band Gaps of Solids with Interacting-Bath Dynamical Embedding Theory

Periodic equation-of-motion coupled-cluster theory with single and double excitations (EOM-CCSD) has shown promise for quantitative calculations of band structures in solids. However, its steep computational scaling has limited calculations to relatively coarse $k$-point meshes, leading to sizable finite-size errors and discrepant estimates of thermodynamic-limit band gaps in recent benchmarks. In this work, we revisit EOM-CCSD band gaps for ten semiconductors and insulators using interacting-bath dynamical embedding theory (ibDET), a systematically improvable Green's function embedding framework that enables dense Brillouin-zone sampling at modest computational cost. By pushing the $k$-point sampling up to $10\times10\times10$, well beyond the system sizes accessible in canonical periodic EOM-CCSD calculations, we significantly reduce finite-size errors and obtain stable thermodynamic-limit extrapolations. We further compare $G_0W_0$@PBE, $G_0W_0$@HF, and EOM-CCSD on an equal footing using the same numerical settings in PySCF. We find that EOM-CCSD yields a mean absolute error of 0.32 eV relative to experimental band gaps for a test set of ten semiconductors and insulators, lower than that of $G_0W_0$@PBE. For ZnO, EOM-CCSD also accurately describes the Zn $3d$-band binding energy, despite overestimating the band gap. These results demonstrate that ibDET offers a practical route to high-accuracy many-body electronic structure calculations in periodic systems.

cond-mat.mtrl-sci

Interpolative Separable Density-Fitting for Transcorrelated Hamiltonians

The transcorrelated (TC) method dramatically accelerates the convergence of correlated calculations toward the complete-basis-set (CBS) limit by folding a Jastrow correlator into the Hamiltonian via a similarity transformation, incorporating the electron--electron cusp into the effective interaction. We make the TC framework practical for large systems and flexible, multi-center correlators by compressing the grid-evaluated TC integrals with the interpolative separable density-fitting (ISDF) approximation, combined with the effective two-body (xTC) treatment of the three-body operator. This low-rank representation reduces storage and integration costs by orders of magnitude, and a multi-GPU implementation with automatic differentiation of the correlator makes the construction routine for large basis sets. We demonstrate the resulting ISDF-xTC-CCSD method on the linear hydrogen chain, reaching the joint thermodynamic and CBS limits with basis sets up to cc-pV5Z in agreement with state-of-the-art many-body references to within about 1~mHa/atom, and on the benzene ground-state energy with up to 1200 orbitals (cc-pCV5Z), where the method attains state-of-the-art accuracy at the coupled cluster singles and doubles level and its CBS extrapolation is markedly more robust than that of conventional coupled-cluster methods.

physics.chem-ph

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI

Embodied AI is moving from isolated perception or action modules toward physical agents that understand, plan under goals, act through robot bodies, monitor progress, and improve from experience. Existing systems address this loop only in parts: end-to-end policies generate actions but often lack spatial reasoning, planning, and execution assessment, while robot-agent systems orchestrate tools or specialists but do not learn a shared representation. This fragmentation limits general Physical Agentic AI. We present ACE-Brain-0.5, a unified embodied foundation model that organizes robot intelligence into five coupled functions: spatial perception, decision making, embodied interaction, self-monitoring, and self-improvement. Built on ACE-Brain-0, which established spatial intelligence as a shared scaffold across robot platforms, ACE-Brain-0.5 extends an understanding-centric model into a closed-loop foundation model. A single 8B backbone instantiates the first four functions: grounding objects and affordances, reasoning over 3D and egocentric spatial relations, decomposing instructions into subgoals, generating navigation and manipulation actions, and estimating progress for verification and recovery. To unify these capabilities without cross-task interference, we introduce SSR+, which extends Scaffold-Specialize-Reconcile with a Reactivate stage after task-vector merging. The fifth function, self-improvement, is realized by a companion framework that updates external execution state, including task schemas, spatial memory, and failure-recovery cases, from rollouts. Across fifteen benchmarks, ACE-Brain-0.5 improves over ACE-Brain-0 on 14 of 18 spatial perception and grounding benchmarks, achieves competitive navigation and manipulation performance, and provides strong progress estimation in ID and OOD settings. Together, these results mark an early step toward general Physical Agentic AI.

cs.RO

A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method

Molecular function is largely determined by structure. Accurately aligning molecular structure with natural language is therefore essential for enabling large language models (LLMs) to reason about downstream chemical tasks. However, the substantial cost of human annotation makes it infeasible to construct large-scale, high-quality datasets of structure-grounded descriptions. In this work, we propose a fully automated annotation framework for generating precise molecular descriptions that preserve complete structural details at scale. Our approach builds upon and extends a rule-based chemical nomenclature parser to interpret IUPAC names and construct enriched, structural XML metadata that explicitly encodes molecular structure. This metadata is then used to guide LLMs in producing accurate natural-language descriptions. Using this framework, we curate a large-scale dataset of approximately $163$k molecule--description pairs. A rigorous validation protocol combining LLM-based and expert human evaluation on a subset of $2,000$ molecules demonstrates a high description precision of $98.6$%. The proposed annotation framework is readily beneficial to broader chemical tasks that rely on structural descriptions, with the resulting dataset providing a reliable foundation for molecule--language alignment. The source code and dataset are hosted at https://github.com/TheLuoFengLab/MolLangData and https://huggingface.co/datasets/ChemFM/MolLangData, respectively.

cs.CL

Boosting Text-Driven Video Segmentation via Geometry-Aware Distillation

Text-driven Referring Video Object Segmentation (RVOS) aims to locate and segment target objects in videos given natural language. However, existing models are typically trained on 2D image or video datasets with naive segmentation losses, which overlooks the geometric consistency across frames and leads to weak spatial understanding. In this paper, we propose Geometry-enhanced Language-guided Video segmentation (GeoLaV), a two-stage framework that distills 3D geometric knowledge from images to enhance text-driven video segmentation. In the first stage, we perform monocular geometry pretraining with monocular novel-view synthesis, enabling the model to acquire geometry-consistent visual representations via spatial alignment on large-scale single-image datasets. In the second stage, we introduce geometry-aware distillation and fine-tune the model on video segmentation datasets, transferring 3D structural knowledge from a general 3D prior model. This process reinforces 3D awareness and improves both spatiotemporal coherence and language grounding in segmentation. Extensive experiments show that our method using only image segmentation data already provides notable zero-shot generalization in RVOS. When combined with geometry-aware distillation for fine-tuning on videos, our method achieves state-of-the-art performance across multiple RVOS benchmarks. The code is available at https://github.com/Tony1882880/GeoLaV.

cs.CV

Low-Scaling Many-Body Green's Function Calculations for Molecular Systems via Interacting-Bath Dynamical Embedding Theory

We present a molecular extension of our recently proposed Green's function embedding method, interacting-bath dynamical embedding theory (ibDET), for computing charged excitation energies at the $GW$ and EOM-CCSD levels. Starting from atom-centered impurities, we construct bath representations that capture the frequency-dependent entanglement between the impurity and its environment and can be systematically improved via the construction of cluster-specific natural orbitals. Utilizing a $GW$ or coupled-cluster Green's function solver, the self-energy of the full system is assembled from all embedding problems to obtain the interacting Green's function. We show that ibDET provides accurate spectral properties with much reduced cost for a broad range of systems, including conjugated molecules and nanoclusters. Compared with full-system results, the errors in the predicted ionization potentials and electron affinities are around 0.1 eV or smaller, while each embedding problem includes only a small fraction of the total orbital space. This work provides an efficient and scalable framework for computing spectral properties of molecular systems.

physics.chem-ph

Transferable Machine Learning of Electronic Hamiltonians with Superposition-of-Atomic-Potentials Features

Machine learning (ML) of electronic Hamiltonians offers a unified route to electronic wave functions and physical observables. We introduce a Hamiltonian learning framework built on electronic features derived from the superposition-of-atomic-potentials (SAP) approximation, an efficient self-consistent-field initial guess that captures essential electron-electron screening. SAP quantities define a symmetry-adapted intrinsic atomic orbital learning basis and provide physics-informed inputs to an orbital-based graph neural network that predicts converged Kohn-Sham Fock matrices. To extend the approach to larger basis sets, we further develop a downfolding scheme that predicts large-basis electronic structure from minimal-basis features. On the QM9 dataset, the model accurately reproduces frontier and core orbital energies, dipole moments, and the full density of states. For organic charge-transport materials, it yields accurate intermolecular transfer integrals for benzene, tetracyanoquinodimethane (TCNQ), and tetrathiafulvalene (TTF) dimers, and transfers to unseen substituted-benzene heterodimers with a mean absolute error of 4.8 meV. These results establish SAP-based ML of electronic Hamiltonians as a transferable and scalable tool for high-throughput electronic-structure prediction.

physics.chem-ph

Integrating Weather Foundation Model and Satellite to Enable Fine-Grained Solar Irradiance Forecasting

Accurate day-ahead solar irradiance forecasting is essential for integrating solar energy into the power grid. However, it remains challenging due to the pronounced diurnal cycle and inherently complex cloud dynamics. Current methods either lack fine-scale resolution (e.g., numerical weather prediction, weather foundation models) or degrade at longer lead times (e.g., satellite extrapolation). We propose Baguan-solar, a two-stage multimodal framework that fuses forecasts from Baguan, a global weather foundation model, with high-resolution geostationary satellite imagery to produce 24-hour irradiance forecasts at kilometer scale. Its decoupled two-stage design first forecasts day-night continuous intermediates (e.g., cloud cover) and then infers irradiance, while its modality fusion jointly preserves fine-scale cloud structures from satellite and large-scale constraints from Baguan forecasts. Evaluated over East Asia using CLDAS as ground truth, Baguan-solar outperforms strong baselines (including ECMWF IFS, vanilla Baguan, and SolarSeer), reducing RMSE by 16.08% and better resolving cloud-induced transients. An operational deployment of Baguan-solar has supported solar power forecasting in an eastern province in China, since July 2025. Our code is accessible at https://github.com/DAMO-DI-ML/Baguan-solar.git.

cs.LG

The Python Simulations of Chemistry Framework: 10 years of an open-source quantum chemistry project

Over the past decade, the Python-based Simulations of Chemistry Framework (PySCF) has developed into a widely used open-source platform for electronic structure theory and quantum chemical method development. This article reviews the major advances since the previous overview in 2020, covering new modules and methodology, infrastructure changes, and performance benchmarks.

physics.chem-ph

MolLangBench: A Comprehensive Benchmark for Language-Prompted Molecular Structure Recognition, Editing, and Generation

Precise recognition, editing, and generation of molecules are essential prerequisites for both chemists and AI systems tackling various chemical tasks. We present MolLangBench, a comprehensive benchmark designed to evaluate fundamental molecule-language interface tasks: language-prompted molecular structure recognition, editing, and generation. To ensure high-quality, unambiguous, and deterministic outputs, we construct the recognition tasks using automated cheminformatics tools, and curate editing and generation tasks through rigorous expert annotation and validation. MolLangBench supports the evaluation of models that interface language with different molecular representations, including linear strings, molecular images, and molecular graphs. Evaluations of state-of-the-art models reveal significant limitations: the strongest model (GPT-5) achieves $86.2\%$ and $85.5\%$ accuracy on recognition and editing tasks, which are intuitively simple for humans, and performs even worse on the generation task, reaching only $43.0\%$ accuracy. These results highlight the shortcomings of current AI systems in handling even preliminary molecular recognition and manipulation tasks. We hope MolLangBench will catalyze further research toward more effective and reliable AI systems for chemical applications.The dataset and code can be accessed at https://huggingface.co/datasets/ChemFM/MolLangBench and https://github.com/TheLuoFengLab/MolLangBench, respectively.

cs.CL

LibppRPA: An Open-Source Library for Particle-Particle Random Phase Approximation

The accurate description of electron correlation and excitation energies remains a fundamental challenge in quantum chemistry. The particle-particle random phase approximation (ppRPA) has emerged as a promising method for capturing a broad range of excited-state properties. However, the implementation of ppRPA has been largely limited to in-house software, restricting its accessibility and usability. In this work, we present LibppRPA, an open-source and lightweight Python library designed for efficient and flexible ppRPA calculations of (1) electronic excitation energy and its associated analytical gradients and (2) the ground state correlation energy, and its associated analytical gradients. LibppRPA enables seamless integration with existing quantum chemistry packages, such as PySCF, by utilizing occupation numbers, molecular orbital coefficients, and three-center electron repulsion integrals. We implement both direct diagonalization and the iterative Davidson algorithm for solving the ppRPA equations, as well as active-space approximations, allowing users to balance accuracy and computational efficiency. We demonstrate the performance of LibppRPA through benchmark calculations on singlet-triplet gaps, double excitations, charge-transfer excitations, and valence/Rydberg excitations, showcasing its reliability across diverse molecular systems. The library provides a robust platform for studying electronic excitations and offers new opportunities for future developments in electronic structure theory.

physics.chem-ph

ChemFM as a Scaling Law Guided Foundation Model Pre-trained on Informative Chemicals

Traditional AI methods often rely on task-specific model designs and training, which constrain both the scalability of model size and generalization across different tasks. Here, we introduce ChemFM, a large foundation model specifically developed for chemicals. By conducting a series of scaling experiments, we identify UniChem as the informative molecular database for pre-training the foundation model. ChemFM comprises 3 billion parameters and is pre-trained on 178 million molecules using self-supervised causal language modeling to extract generalizable molecular representations. This model can be adapted to diverse downstream chemical applications using either full-parameter or parameter-efficient fine-tuning methods. ChemFM consistently outperforms state-of-the-art task-specific AI models across all tested tasks. Notably, it achieves up to 67.48% performance improvement across 34 property prediction benchmarks, up to 33.80% reduction in mean average deviation between conditioned and actual properties of generated molecules in conditional molecular generation tasks, and up to 3.7% top-1 accuracy improvement across 4 reaction prediction datasets. Moreover, ChemFM demonstrates its superior performance in predicting antibiotic activity and cytotoxicity, highlighting its potential to advance the discovery of novel antibiotics. Furthermore, we demonstrate that, as a foundation model, ChemFM exhibits strong data efficiency, requiring significantly fewer labeled training samples to achieve state-of-the-art performance. We anticipate that ChemFM will significantly advance chemistry research by providing a foundation model capable of effectively generalizing across a broad range of tasks with minimal additional training.

cs.CE

Towards an exact electronic quantum many-body treatment of Kondo correlation in magnetic impurities

The Kondo effect is a prototypical quantum phenomenon arising from the interaction between localized electrons in a magnetic impurity and itinerant electrons in a metallic host. Although it has served as the testing ground for quantum many-body methods for decades, the precise description of Kondo physics with material specificity remains challenging. Here, we present a systematic ab initio approach to converge towards an exact zero-temperature electronic treatment of Kondo correlations. Across a series of 3d transition metals, we extract Kondo temperatures matching the subtle experimental trends, with an accuracy exceeding that of standard models. We further obtain microscopic insight into the origin of these trends. More broadly, we demonstrate the possibility to start from fully ab initio many-body simulations and push towards the realm of converged predictions.

cond-mat.str-el

Unified Deep Learning Framework for Many-Body Quantum Chemistry via Green's Functions

Quantum many-body methods provide a systematic route to computing electronic properties of molecules and materials, but high computational costs restrict their use in large-scale applications. Due to the complexity in many-electron wavefunctions, machine learning models capable of capturing fundamental many-body physics remain limited. Here, we present a deep learning framework targeting the many-body Green's function, which unifies predictions of electronic properties in ground and excited states, while offering physical insights into many-electron correlation effects. By learning the $GW$ or coupled-cluster self-energy from mean-field features, our graph neural network achieves competitive performance in predicting one- and two-particle excitations and quantities derivable from one-particle density matrix. We demonstrate its high data efficiency and good transferability across chemical species, system sizes, molecular conformations, and correlation strengths in bond breaking, through multiple molecular and nanomaterial benchmarks. This work opens up opportunities for utilizing machine learning to solve many-electron problems.

physics.chem-ph

Generalized Contrastive Learning for Multi-Modal Retrieval and Ranking

Contrastive learning has gained widespread adoption for retrieval tasks due to its minimal requirement for manual annotations. However, popular training frameworks typically learn from binary (positive/negative) relevance, making them ineffective at incorporating desired rankings. As a result, the poor ranking performance of these models forces systems to employ a re-ranker, which increases complexity, maintenance effort and inference time. To address this, we introduce Generalized Contrastive Learning (GCL), a training framework designed to learn from continuous ranking scores beyond binary relevance. GCL encodes both relevance and ranking information into a unified embedding space by applying ranking scores to the loss function. This enables a single-stage retrieval system. In addition, during our research, we identified a lack of public multi-modal datasets that benchmark both retrieval and ranking capabilities. To facilitate this and future research for ranked retrieval, we curated a large-scale MarqoGS-10M dataset using GPT-4 and Google Shopping, providing ranking scores for each of the 10 million query-document pairs. Our results show that GCL achieves a 29.3% increase in NDCG@10 for in-domain evaluations and 6.0% to 10.0% increases for cold-start evaluations compared to the finetuned CLIP baseline with MarqoGS-10M. Additionally, we evaluated GCL offline on a proprietary user interaction data. GCL shows an 11.2% gain for in-domain evaluations. The dataset and the method are available at: https://github.com/marqo-ai/GCL.

cs.IR

Energy-Specific Bethe-Salpeter Equation Implementation for Efficient Optical Spectrum Calculations

We present an energy-specific Bethe-Salpeter equation (BSE) implementation for efficient core and valence optical spectrum calculations. In energy-specific BSE, high-lying excitation energies are obtained by constructing trial vectors and expanding the subspace targeting excitation energies above the predefined energy threshold in the Davidson algorithm. To calculate optical spectra over a wide energy range, energy-specific BSE can be applied to multiple consecutive small energy windows, where trial vectors for each subsequent energy window are made orthogonal to the subspace of preceding windows to accelerate the convergence of the Davidson algorithm. For seven small molecules, energy-specific BSE combined with $G_0W_0$ provides small errors around 0.8 eV for absolute and relative $K$-edge excitation energies when starting from a hybrid PBEh solution with 45% exact exchange. We further showcase the computational efficiency of this approach by simulating the N $1s$ $K$-edge excitation spectrum of the porphine molecule and the valence optical spectrum of silicon nanoclusters involving 6,000 excited states using $G_0W_0$-BSE. This work expands the applicability of the $GW$-BSE formalism for investigating high-energy excited states of large systems.

cond-mat.mtrl-sci

Video Quality Assessment: A Comprehensive Survey

Video quality assessment (VQA) is an important processing task, aiming at predicting the quality of videos in a manner highly consistent with human judgments of perceived quality. Traditional VQA models based on natural image and/or video statistics, which are inspired both by models of projected images of the real world and by dual models of the human visual system, deliver only limited prediction performances on real-world user-generated content (UGC), as exemplified in recent large-scale VQA databases containing large numbers of diverse video contents crawled from the web. Fortunately, recent advances in deep neural networks and Large Multimodality Models (LMMs) have enabled significant progress in solving this problem, yielding better results than prior handcrafted models. Numerous deep learning-based VQA models have been developed, with progress in this direction driven by the creation of content-diverse, large-scale human-labeled databases that supply ground truth psychometric video quality data. Here, we present a comprehensive survey of recent progress in the development of VQA algorithms and the benchmarking studies and databases that make them possible. We also analyze open research directions on study design and VQA algorithm architectures. Github link: https://github.com/taco-group/Video-Quality-Assessment-A-Comprehensive-Survey.

eess.IV