SearcharxivSearch

arXiv subjects

Yunrui Li

Publications and source records attributed to Yunrui Li.

7 recordsLinked to original sources

Scaling-law-informed neural point processes for earthquake sequence forecasting

Earthquake sequence forecasting requires models that can learn nonlinear history dependence while retaining robust statistical structure. We develop a scaling-law-informed neural marked point process, termed Fusion, that combines neural representations of catalog history with temporal features derived from the Epidemic-Type Aftershock Sequence model and magnitude information derived from the Gutenberg--Richter law. The model separates the magnitude cutoff applied to the input catalog from the fixed target-event threshold, allowing lower-magnitude earthquakes to inform forecasts without changing the target-event set. For the 2016--2017 Amatrice--Visso--Norcia sequence, Fusion achieves the highest target-event temporal likelihood when lower-magnitude events are retained, outperforming both ETAS and a purely neural point-process baseline. Event-wise and cumulative analyses show sustained timing gains through substantial portions of the Visso and Norcia sequences. Across five benchmark catalogs, catalog-specific neural training with a fixed ETAS prior yields the highest temporal likelihood at the minimum evaluated magnitude cutoff. Magnitude likelihood shows no consistent predictive gain beyond the Gutenberg--Richter-based ETAS reference, indicating that the additional information captured by Fusion is primarily temporal. These results show that lower-magnitude catalog histories and empirical scaling-law information complement neural sequence learning for target-event timing.

physics.geo-ph

Multimodal Fusion with Relational Learning for Molecular Property Prediction

Graph based molecular representation learning is essential for accurately predicting molecular properties in drug discovery and materials science; however, it faces significant challenges due to the intricate relationships among molecules and the limited chemical knowledge utilized during training. While contrastive learning is often employed to handle molecular relationships, its reliance on binary metrics is insufficient for capturing the complexity of these interactions. Multimodal fusion has gained attention for property reasoning, but previous work has explored only a limited range of modalities, and the optimal stages for fusing different modalities in molecular property tasks remain underexplored. In this paper, we introduce MMFRL (Multimodal Fusion with Relational Learning for Molecular Property Prediction), a novel framework designed to overcome these limitations. Our method enhances embedding initialization through multimodal pretraining using relational learning. We also conduct a systematic investigation into the impact of modality fusion at different stages such as early, intermediate, and late, highlighting their advantages and shortcomings. Extensive experiments on MoleculeNet benchmarks demonstrate that MMFRL significantly outperforms existing methods. Furthermore, MMFRL enables task-specific optimizations. Additionally, the explainability of MMFRL provides valuable chemical insights, emphasizing its potential to enhance real-world drug discovery applications.

cs.CE

2DNMRGym: An Annotated Experimental Dataset for Atom-Level Molecular Representation Learning in 2D NMR via Surrogate Supervision

Two-dimensional (2D) Nuclear Magnetic Resonance (NMR) spectroscopy, particularly Heteronuclear Single Quantum Coherence (HSQC) spectroscopy, plays a critical role in elucidating molecular structures, interactions, and electronic properties. However, accurately interpreting 2D NMR data remains labor-intensive and error-prone, requiring highly trained domain experts, especially for complex molecules. Machine Learning (ML) holds significant potential in 2D NMR analysis by learning molecular representations and recognizing complex patterns from data. However, progress has been limited by the lack of large-scale and high-quality annotated datasets. In this work, we introduce 2DNMRGym, the first annotated experimental dataset designed for ML-based molecular representation learning in 2D NMR. It includes over 22,000 HSQC spectra, along with the corresponding molecular graphs and SMILES strings. Uniquely, 2DNMRGym adopts a surrogate supervision setup: models are trained using algorithm-generated annotations derived from a previously validated method and evaluated on a held-out set of human-annotated gold-standard labels. This enables rigorous assessment of a model's ability to generalize from imperfect supervision to expert-level interpretation. We provide benchmark results using a series of 2D and 3D GNN and GNN transformer models, establishing a strong foundation for future work. 2DNMRGym supports scalable model training and introduces a chemically meaningful benchmark for evaluating atom-level molecular representations in NMR-guided structural tasks. Our data and code is open-source and available on Huggingface and Github.

cs.LG

Advancing Drug Discovery with Enhanced Chemical Understanding via Asymmetric Contrastive Multimodal Learning

The versatility of multimodal deep learning holds tremendous promise for advancing scientific research and practical applications. As this field continues to evolve, the collective power of cross-modal analysis promises to drive transformative innovations, opening new frontiers in chemical understanding and drug discovery. Hence, we introduce Asymmetric Contrastive Multimodal Learning (ACML), a specifically designed approach to enhance molecular understanding and accelerate advancements in drug discovery. ACML harnesses the power of effective asymmetric contrastive learning to seamlessly transfer information from various chemical modalities to molecular graph representations. By combining pre-trained chemical unimodal encoders and a shallow-designed graph encoder with 5 layers, ACML facilitates the assimilation of coordinated chemical semantics from different modalities, leading to comprehensive representation learning with efficient training. We demonstrate the effectiveness of this framework through large-scale cross-modality retrieval and isomer discrimination tasks. Additionally, ACML enhances interpretability by revealing chemical semantics in graph presentations and bolsters the expressive power of graph neural networks, as evidenced by improved performance in molecular property prediction tasks from MoleculeNet and Therapeutics Data Commons (TDC). Ultimately, ACML exemplifies its potential to revolutionize molecular representational learning, offering deeper insights into the chemical semantics of diverse modalities and paving the way for groundbreaking advancements in chemical research and drug discovery.

cs.LG

TransPeakNet: Solvent-Aware 2D NMR Prediction via Multi-Task Pre-Training and Unsupervised Learning

Nuclear Magnetic Resonance (NMR) spectroscopy is essential for revealing molecular structure, electronic environment, and dynamics. Accurate NMR shift prediction allows researchers to validate structures by comparing predicted and observed shifts. While Machine Learning (ML) has improved one-dimensional (1D) NMR shift prediction, predicting 2D NMR remains challenging due to limited annotated data. To address this, we introduce an unsupervised training framework for predicting cross-peaks in 2D NMR, specifically Heteronuclear Single Quantum Coherence (HSQC).Our approach pretrains an ML model on an annotated 1D dataset of 1H and 13C shifts, then finetunes it in an unsupervised manner using unlabeled HSQC data, which simultaneously generates cross-peak annotations. Our model also adjusts for solvent effects. Evaluation on 479 expert-annotated HSQC spectra demonstrates our model's superiority over traditional methods (ChemDraw and Mestrenova), achieving Mean Absolute Errors (MAEs) of 2.05 ppm and 0.165 ppm for 13C shifts and 1H shifts respectively. Our algorithmic annotations show a 95.21% concordance with experts' assignments, underscoring the approach's potential for structural elucidation in fields like organic chemistry, pharmaceuticals, and natural products.

cs.LG

Deep-learning Optical Flow Outperforms PIV in Obtaining Velocity Fields from Active Nematics

Deep learning-based optical flow (DLOF) extracts features in adjacent video frames with deep convolutional neural networks. It uses those features to estimate the inter-frame motions of objects at the pixel level. In this article, we evaluate the ability of optical flow to quantify the spontaneous flows of MT-based active nematics under different labeling conditions. We compare DLOF against the commonly used technique, particle imaging velocimetry (PIV). We obtain flow velocity ground truths either by performing semi-automated particle tracking on samples with sparsely labeled filaments, or from passive tracer beads. We find that DLOF produces significantly more accurate velocity fields than PIV for densely labeled samples. We show that the breakdown of PIV arises because the algorithm cannot reliably distinguish contrast variations at high densities, particularly in directions parallel to the nematic director. DLOF overcomes this limitation. For sparsely labeled samples, DLOF and PIV produce results with similar accuracy, but DLOF gives higher-resolution fields. Our work establishes DLOF as a versatile tool for measuring fluid flows in a broad class of active, soft, and biophysical systems.

cond-mat.soft

A Machine Learning Approach to Robustly Determine Director Fields and Analyze Defects in Active Nematics

Active nematics are dense systems of rodlike particles that consume energy to drive motion at the level of the individual particles. They exist in natural systems like biological tissues and artificial materials such as suspensions of self-propelled colloidal particles or synthetic microswimmers. Active nematics have attracted significant attention in recent years due to their spectacular nonequilibrium collective spatiotemporal dynamics, which may enable applications in fields such as robotics, drug delivery, and materials science. The director field, which measures the direction and degree of alignment of the local nematic orientation, is a crucial characteristic of active nematic and is essential for studying topological defects. However, determining the director field is a significant challenge in many experimental systems. Although director fields can be derived from images of active nematics using traditional imaging processing methods, the accuracy of such methods are highly sensitive to the settings of the algorithms. These settings must be tuned from image-to-image due to experimental noise, intrinsic noise of the imaging technology, and perturbations caused by changes in experimental conditions. This sensitivity currently limits automatic analysis of active nematics. To address this, we developed a machine learning model for extracting reliable director fields from raw experimental images, which enables accurate analysis of topological defects. Application of the algorithm to experimental data demonstrates that the approach is robust and highly generalizable to experimental settings that are different from those in the training data. It could be a promising tool for investigating active nematics and may be generalized to other active matter systems.

cond-mat.soft