SearcharxivSearch

arXiv subjects

Lijiang Yang

Publications and source records attributed to Lijiang Yang.

11 recordsLinked to original sources

Multi-objective fluorescent molecule design with a data-physics dual-driven generative framework

Designing fluorescent small molecules with tailored optical and physicochemical properties requires navigating vast, underexplored chemical space while satisfying multiple objectives and constraints. Conventional generate-score-screen approaches become impractical under such realistic design specifications, owing to their low search efficiency, unreliable generalizability of machine-learning prediction, and the prohibitive cost of quantum chemical calculation. Here we present LUMOS, a data-and-physics driven framework for inverse design of fluorescent molecules. LUMOS couples generator and predictor within a shared latent representation, enabling direct specification-to-molecule design and efficient exploration. Moreover, LUMOS combines neural networks with a fast time-dependent density functional theory (TD-DFT) calculation workflow to build a suite of complementary predictors spanning different trade-offs in speed, accuracy, and generalizability, enabling reliable property prediction across diverse scenarios. Finally, LUMOS employs a property-guided diffusion model integrated with multi-objective evolutionary algorithms, enabling de novo design and molecular optimization under multiple objectives and constraints. Across comprehensive benchmarks, LUMOS consistently outperforms baseline models in terms of accuracy, generalizability and physical plausibility for fluorescence property prediction, and demonstrates superior performance in multi-objective scaffold- and fragment-level molecular optimization. Further validation using TD-DFT and molecular dynamics (MD) simulations demonstrates that LUMOS can generate valid fluorophores that meet various target specifications. Overall, these results establish LUMOS as a data-physics dual-driven framework for general fluorophore inverse design.

cs.LG

TRIDS: AI-native molecular docking framework for accelerating high-throughput virtual screening with physically valid binding poses

Molecular docking is a cornerstone of drug discovery for unveiling the mechanism of ligand-receptor interactions. With the recent advances of deep learning (DL), AI-powered molecular docking methods have achieved higher accuracy for binding pose prediction and virtual screening compared with classical physics-based methods. However, there is still a scarcity of approaches to strike a balance among accuracy, computational efficiency, and rigorous physical validity of the output conformations. In the previous two versions of DSDP, we demonstrated the effectiveness of guiding conformation sampling with the gradient of an analytic scoring function. As the third version, TRIDS was devised as an AI-native docking framework that expand the similar strategy to unify conformation sampling and docking processes with DL-based model for improving accuracy of docking and screening. Furthermore, it is tailored for seamless cooperation of AI and physics to guarantee the physical validity of predicted binding poses. Being user-friendly, TRIDS predicts the binding site, parses multiple file formats, and supports Python programming and PyMOL graphical interaction. It improves docking accuracy and passes physical validation with high computational efficiency, i.e. a single docking task is done in a fraction of second while maintaining a highly lightweight GPU memory footprint of merely hundreds of megabytes, facilitating high-throughput virtual screening in reality. As a proof of concept, TRIDS allowed us to obtain hit compounds with novel scaffolds for tumor necrosis factor-alpha (TNF{\alpha}) inhibitor through a large-scale virtual screening.

physics.chem-ph

Large Language Models as AI Agents for Digital Atoms and Molecules: Catalyzing a New Era in Computational Biophysics

In computational biophysics, where molecular data is expanding rapidly and system complexity is increasing exponentially, large language models (LLMs) and agent-based systems are fundamentally reshaping the field. This perspective article examines the recent advances at the intersection of LLMs, intelligent agents, and scientific computation, with a focus on biophysical computation. Building on these advancements, we introduce ADAM (Agent for Digital Atoms and Molecules), an innovative multi-agent LLM-based framework. ADAM employs cutting-edge AI architectures to reshape scientific workflows through a modular design. It adopts a hybrid neural-symbolic architecture that combines LLM-driven semantic tools with deterministic symbolic computations. Moreover, its ADAM Tool Protocol (ATP) enables asynchronous, database-centric tool orchestration, fostering community-driven extensibility. Despite the significant progress made, ongoing challenges call for further efforts in establishing benchmarking standards, optimizing foundational models and agents, building an open collaborative ecosystem and developing personalized memory modules. ADAM is accessible at https://sidereus-ai.com.

physics.comp-ph

A Generalized Nucleation Theory for Ice Crystallization

Despite the simplicity of the water molecule, the kinetics of ice nucleation under natural conditions can be complex. We investigated spontaneously grown ice nuclei using all-atom molecular dynamics simulations and found significant differences between the kinetics of ice formation through spontaneously formed and ideal nuclei. Since classical nucleation theory can only provide a good description of ice nucleation in ideal conditions, we propose a generalized nucleation theory that can better characterize the kinetics of ice crystal nucleation in general conditions. This study provides an explanation on why previous experimental and computational studies have yielded widely varying critical nucleation sizes.

cond-mat.soft

Enhanced sampling reveals the main pathway of an organic cascade reaction

Normal molecular dynamics simulations are usually unable to simulate chemical reactions due to the low probability of forming the transition state. Therefore, enhanced sampling methods are implemented to accelerate the occurrence of chemical reactions. In this investigation, we present an application of metadynamics in simulating an organic multi-step cascade reaction. The analysis of the reaction trajectory reveals the barrier heights of both forward and reverse reactions. We also present a discussion of the advantages and disadvantages of generating the reactive pathway using molecular dynamics and the intrinsic reaction coordinate (IRC) algorithm.

physics.chem-ph

Unsupervisedly Prompting AlphaFold2 for Few-Shot Learning of Accurate Folding Landscape and Protein Structure Prediction

Data-driven predictive methods which can efficiently and accurately transform protein sequences into biologically active structures are highly valuable for scientific research and medical development. Determining accurate folding landscape using co-evolutionary information is fundamental to the success of modern protein structure prediction methods. As the state of the art, AlphaFold2 has dramatically raised the accuracy without performing explicit co-evolutionary analysis. Nevertheless, its performance still shows strong dependence on available sequence homologs. Based on the interrogation on the cause of such dependence, we presented EvoGen, a meta generative model, to remedy the underperformance of AlphaFold2 for poor MSA targets. By prompting the model with calibrated or virtually generated homologue sequences, EvoGen helps AlphaFold2 fold accurately in low-data regime and even achieve encouraging performance with single-sequence predictions. Being able to make accurate predictions with few-shot MSA not only generalizes AlphaFold2 better for orphan sequences, but also democratizes its use for high-throughput applications. Besides, EvoGen combined with AlphaFold2 yields a probabilistic structure generation method which could explore alternative conformations of protein sequences, and the task-aware differentiable algorithm for sequence generation will benefit other related tasks including protein design.

cs.LG

PSP: Million-level Protein Sequence Dataset for Protein Structure Prediction

Proteins are essential component of human life and their structures are important for function and mechanism analysis. Recent work has shown the potential of AI-driven methods for protein structure prediction. However, the development of new models is restricted by the lack of dataset and benchmark training procedure. To the best of our knowledge, the existing open source datasets are far less to satisfy the needs of modern protein sequence-structure related research. To solve this problem, we present the first million-level protein structure prediction dataset with high coverage and diversity, named as PSP. This dataset consists of 570k true structure sequences (10TB) and 745k complementary distillation sequences (15TB). We provide in addition the benchmark training procedure for SOTA protein structure prediction model on this dataset. We validate the utility of this dataset for training by participating CAMEO contest in which our model won the first place. We hope our PSP dataset together with the training benchmark can enable a broader community of AI/biology researchers for AI-driven protein related research.

q-bio.BM

Atomistic View of Homogeneous Nucleation of Water into Polymorphic Ices

Water is one of the most abundant substances on Earth, and ice, i.e., solid water, has more than 18 known phases. Normally ice in nature exists only as Ice Ih, Ice Ic, or a stacking disordered mixture of both. Although many theoretical efforts have been devoted to understanding the thermodynamics of different ice phases at ambient temperature and pressure, there still remains many puzzles. We simulated the reversible transitions between water and different ice phases by performing full atom molecular dynamics simulations. Using the enhanced sampling method MetaITS with the two selected X-ray diffraction peak intensities as collective variables, the ternary phase diagrams of liquid water, ice Ih, ice Ic at multiple were obtained. We also present a simple physical model which successfully explains the thermodynamic stability of ice. Our results agree with experiments and leads to a deeper understanding of the ice nucleation mechanism.

cond-mat.stat-mech

Deep Reinforcement Learning of Transition States

Combining reinforcement learning (RL) and molecular dynamics (MD) simulations, we propose a machine-learning approach (RL$^‡$) to automatically unravel chemical reaction mechanisms. In RL$^‡$, locating the transition state of a chemical reaction is formulated as a game, where a virtual player is trained to shoot simulation trajectories connecting the reactant and product. The player utilizes two functions, one for value estimation and the other for policy making, to iteratively improve the chance of winning this game. We can directly interpret the reaction mechanism according to the value function. Meanwhile, the policy function enables efficient sampling of the transition paths, which can be further used to analyze the reaction dynamics and kinetics. Through multiple experiments, we show that RL‡ can be trained tabula rasa hence allows us to reveal chemical reaction mechanisms with minimal subjective biases.

physics.chem-ph

A Perspective on Deep Learning for Molecular Modeling and Simulations

Deep learning is transforming many areas in science, and it has great potential in modeling molecular systems. However, unlike the mature deployment of deep learning in computer vision and natural language processing, its development in molecular modeling and simulations is still at an early stage, largely because the inductive biases of molecules are completely different from those of images or texts. Footed on these differences, we first reviewed the limitations of traditional deep learning models from the perspective of molecular physics, and wrapped up some relevant technical advancement at the interface between molecular modeling and deep learning. We do not focus merely on the ever more complex neural network models, instead, we emphasize the theories and ideas behind modern deep learning. We hope that transacting these ideas into molecular modeling will create new opportunities. For this purpose, we summarized several representative applications, ranging from supervised to unsupervised and reinforcement learning, and discussed their connections with the emerging trends in deep learning. Finally, we outlook promising directions which may help address the existing issues in the current framework of deep molecular modeling.

physics.comp-ph

Combine Umbrella Sampling with Integrated Tempering Method for Efficient and Accurate Calculation of Free Energy Changes of Complex Energy Surface

Umbrella sampling is an efficient method for the calculation of free energy changes of a system along well-defined reaction coordinates. However, when multiple parallel channels along the reaction coordinate or hidden barriers in directions perpendicular to the reaction coordinate exist, it is difficult for conventional umbrella sampling methods to generate sufficient sampling within limited simulation time. Here we propose an efficient approach to combine umbrella sampling with the integrated tempering sampling method. The umbrella sampling method is applied to conformational degrees of freedom which possess significant barriers and are chemically more relevant. The integrated tempering sampling method is employed to facilitate the sampling of other degrees of freedom in which statistically non-negligible barriers may exist. The combined method is applied to two model systems and show significantly improved sampling efficiencies as compared to standalone conventional umbrella sampling or integrated tempering sampling approaches. Therefore, the combined approach will become a very efficient method in the simulation of biomolecular processes which often involve sampling of complex rugged energy landscapes.

stat.ME