SearcharxivSearch

arXiv subjects

Yanming Wang

Publications and source records attributed to Yanming Wang.

At least 19 recordsLinked to original sources

EvoMD-LLM: Learning the Language of Species Evolution in Reactive Molecular Dynamics

While large language models (LLMs) excel at static scientific reasoning, they struggle to model the temporal structure of dynamic physical processes. We present EvoMD-LLM (Evolutionary Molecular Dynamics Large Language Model), a framework that reformulates species-level molecular dynamics as a symbolic temporal language modeling problem. Reactive MD trajectories are discretized into sequences of molecular events, where each token represents a chemical species augmented with its persistence duration, enabling standard autoregressive LLMs to learn compositional evolution over time through efficient fine-tuning. A key component of EvoMD-LLM is temporal scaffolding, which treats event duration as an explicit linguistic token and serves as a structured inductive bias, significantly reducing invalid or hallucinated molecular outputs compared to conventional sequence modeling approaches. We evaluate EvoMD-LLM on multiple temporal prediction tasks, achieving up to 66.14% accuracy and consistently outperforming sequential neural networks and language-based baselines. Beyond quantitative improvements, we qualitatively observe that the model is capable of generating interpretations for its own predictions by incorporating relevant chemical knowledge, even though it was not explicitly supervised with paired trajectory-explanation data. These results demonstrate that symbolic temporal language modeling provides an effective framework for grounding LLMs in dynamic physical simulations.

cs.AI

PolyCrysDiff: Controllable Generation of Three-Dimensional Computable Polycrystalline Material Structures

The three-dimensional (3D) microstructures of polycrystalline materials exert a critical influence on their mechanical and physical properties. Realistic, controllable construction of these microstructures is a key step toward elucidating structure-property relationships, yet remains a formidable challenge. Herein, we propose PolyCrysDiff, a framework based on conditional latent diffusion that enables the end-to-end generation of computable 3D polycrystalline microstructures. Comprehensive qualitative and quantitative evaluations demonstrate that PolyCrysDiff faithfully reproduces target grain morphologies, orientation distributions, and 3D spatial correlations, while achieving an $R^2$ over 0.972 on grain attributes (e.g., size and sphericity) control, thereby outperforming mainstream approaches such as Markov random field (MRF)- and convolutional neural network (CNN)-based methods. The computability and physical validity of the generated microstructures are verified through a series of crystal plasticity finite element method (CPFEM) simulations. Leveraging PolyCrysDiff's controllable generative capability, we systematically elucidate how grain-level microstructural characteristics affect the mechanical properties of polycrystalline materials. This development is expected to pave a key step toward accelerated, data-driven optimization and design of polycrystalline materials.

cs.CV

Physics Informed Generative AI Enabling Labour Free Segmentation For Microscopy Analysis

Semantic segmentation of microscopy images is a critical task for high-throughput materials characterisation, yet its automation is severely constrained by the prohibitive cost, subjectivity, and scarcity of expert-annotated data. While physics-based simulations offer a scalable alternative to manual labelling, models trained on such data historically fail to generalise due to a significant domain gap, lacking the complex textures, noise patterns, and imaging artefacts inherent to experimental data. This paper introduces a novel framework for labour-free segmentation that successfully bridges this simulation-to-reality gap. Our pipeline leverages phase-field simulations to generate an abundant source of microstructural morphologies with perfect, intrinsically-derived ground-truth masks. We then employ a Cycle-Consistent Generative Adversarial Network (CycleGAN) for unpaired image-to-image translation, transforming the clean simulations into a large-scale dataset of high-fidelity, realistic SEM images. A U-Net model, trained exclusively on this synthetic data, demonstrated remarkable generalisation when deployed on unseen experimental images, achieving a mean Boundary F1-Score of 0.90 and an Intersection over Union (IOU) of 0.88. Comprehensive validation using t-SNE feature-space projection and Shannon entropy analysis confirms that our synthetic images are statistically and featurally indistinguishable from the real data manifold. By completely decoupling model training from manual annotation, our generative framework transforms a data-scarce problem into one of data abundance, providing a robust and fully automated solution to accelerate materials discovery and analysis.

cs.CV

A joint voxel flow-phase field framework for ultra-long microstructure evolution prediction with physical regularization

Phase-field (PF) modeling is a powerful tool for simulating microstructure evolution. To accelerate the simulation of PF models governed by complex PDEs, machine learning methods such as PINNs and ConvLSTM have been introduced. However, current machine-learning-based approaches still suffer from limited flexibility, poor generalization, and short prediction horizons. To address these challenges, we present a joint framework that couples a voxel-flow network (VFN) with PF simulations in an alternating manner for long-horizon prediction of microstructure evolution with substantial computational acceleration. The VFN iteratively predicts future evolution by generating the next snapshot from the previous two snapshots. Periodic PF simulations suppress nonphysical artifacts, reduce accumulated error, and extend the reliable prediction horizon. The VFN was validated using a grain-growth example, and its accuracy outperforms that of similar prediction methods while preserving topological grain details. For an ultra-long grain-growth prediction of 82 frames from 2 input frames, the grain number decreases from 600 to 29 while the NMSE of the average grain area remains 1.64%. The framework also exhibits good generalizability across different PF models. Overall, this joint framework enables rapid, flexible, generalizable, and physically consistent microstructure forecasting from image-based data over ultra-long time scales.

physics.comp-ph

A Large-Language-Model Assisted Automated Scale Bar Detection and Extraction Framework for Scanning Electron Microscopic Images

Microscopic characterizations, such as Scanning Electron Microscopy (SEM), are widely used in scientific research for visualizing and analyzing microstructures. Determining the scale bars is an important first step of accurate SEM analysis; however, currently, it mainly relies on manual operations, which is both time-consuming and prone to errors. To address this issue, we propose a multi-modal and automated scale bar detection and extraction framework that provides concurrent object detection, text detection and text recognition with a Large Language Model (LLM) agent. The proposed framework operates in four phases; i) Automatic Dataset Generation (Auto-DG) model to synthesize a diverse dataset of SEM images ensuring robust training and high generalizability of the model, ii) scale bar object detection, iii) information extraction using a hybrid Optical Character Recognition (OCR) system with DenseNet and Convolutional Recurrent Neural Network (CRNN) based algorithms, iv) an LLM agent to analyze and verify accuracy of the results. The proposed model demonstrates a strong performance in object detection and accurate localization with a precision of 100%, recall of 95.8%, and a mean Average Precision (mAP) of 99.2% at IoU=0.5 and 69.1% at IoU=0.5:0.95. The hybrid OCR system achieved 89% precision, 65% recall, and a 75% F1 score on the Auto-DG dataset, significantly outperforming several mainstream standalone engines, highlighting its reliability for scientific image analysis. The LLM is introduced as a reasoning engine as well as an intelligent assistant that suggests follow-up steps and verifies the results. This automated method powered by an LLM agent significantly enhances the efficiency and accuracy of scale bar detection and extraction in SEM images, providing a valuable tool for microscopic analysis and advancing the field of scientific imaging.

cs.CV

Mass-transport-limited reaction rates and molecular diffusion in the van der Waals gap beneath graphene

The confinement of molecules within the van der Waals (vdW) gap between a two-dimensional 2D material and a catalytic substrate offers a promising route toward the development of molecule-selective catalysts with increased reaction rates. However, identifying the kinetic limitations of such confined reactions remains challenging. Here, we employ an inverted wedding-cake configuration of multilayer graphene on platinum to study the dynamics of graphene etching in the vdW gap by various molecules (O2, H2, and CO), using in situ scanning electron microscopy. Under the experimental conditions explored (up to p = 1.4x10-3 Pa and T = 1000 {\deg}C), the etching reaction rates are limited by mass transport within the confined space. This limitation persists even for CO, despite its anomalously enhanced transport resulting from a significant lifting of the vdW gap. Reactive molecular dynamics simulations further reveal multiple etching pathways for CO, enabled by confinement within the vdW space. Once mass-transport limitations are overcome, the vdW gap acts as an effective nanoreactor, facilitating reaction pathways that would be otherwise inaccessible on a pristine surface without spatial confinement.

cond-mat.mes-hall

Volume-Wise Task fMRI Decoding with Deep Learning:Enhancing Temporal Resolution and Cognitive Function Analysis

In recent years,the application of deep learning in task functional Magnetic Resonance Imaging (tfMRI) decoding has led to significant advancements. However,most studies remain constrained by assumption of temporal stationarity in neural activity,resulting in predominantly block-wise analysis with limited temporal resolution on the order of tens of seconds. This limitation restricts the ability to decode cognitive functions in detail. To address these limitations, this study proposes a deep neural network designed for volume-wise identification of task states within tfMRI data,thereby overcoming the constraints of conventional methods. Evaluated on Human Connectome Project (HCP) motor and gambling tfMRI datasets,the model achieved impressive mean accuracy rates of 94.0% and 79.6%,respectively. These results demonstrate a substantial enhancement in temporal resolution,enabling more detailed exploration of cognitive processes. The study further employs visualization algorithms to investigate dynamic brain mappings during different tasks,marking a significant step forward in deep learning-based frame-level tfMRI decoding. This approach offers new methodologies and tools for examining dynamic changes in brain activities and understanding the underlying cognitive mechanisms.

cs.LG

SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks

With the rapid advancement of Large Language Models (LLMs), the safety of LLMs has been a critical concern requiring precise assessment. Current benchmarks primarily concentrate on single-turn dialogues or a single jailbreak attack method to assess the safety. Additionally, these benchmarks have not taken into account the LLM's capability of identifying and handling unsafe information in detail. To address these issues, we propose a fine-grained benchmark SafeDialBench for evaluating the safety of LLMs across various jailbreak attacks in multi-turn dialogues. Specifically, we design a two-tier hierarchical safety taxonomy that considers 6 safety dimensions and generates more than 4000 multi-turn dialogues in both Chinese and English under 22 dialogue scenarios. We employ 7 jailbreak attack strategies, such as reference attack and purpose reverse, to enhance the dataset quality for dialogue generation. Notably, we construct an innovative assessment framework of LLMs, measuring capabilities in detecting, and handling unsafe information and maintaining consistency when facing jailbreak attacks. Experimental results across 17 LLMs reveal that Yi-34B-Chat and GLM4-9B-Chat demonstrate superior safety performance, while Llama3.1-8B-Instruct and o3-mini exhibit safety vulnerabilities.

cs.CL

Defects in graphite engineered by ion implantation for the self-assembly of gold nanoparticles

Defect engineering in two-dimensional (2D) materials is essential for advancing applications such as gas sensing, single-atom catalysis, and guided nanoparticle self-assembly, enabling the creation of materials with tailored functionalities. This study investigates ion implantation effects on highly ordered pyrolytic graphite (HOPG) surfaces, using scanning tunneling microscopy (STM) and density functional theory (DFT) simulations to identify distinct defect structures. High-energy heavy ions cause inelastic scattering, increasing surface damage, while gold atoms deposited onto defect sites preferentially form atomic clusters. Through focused ion beam techniques, spatially distributed defects were engineered, guiding the self-assembly of nanoparticles. This research highlights the precision of ion irradiation for modifying HOPG surfaces, with significant implications for catalysis, nanotechnology, and the development of functional materials with controlled nanoscale properties.

cond-mat.mtrl-sci

LLMem: Estimating GPU Memory Usage for Fine-Tuning Pre-Trained LLMs

Fine-tuning pre-trained large language models (LLMs) with limited hardware presents challenges due to GPU memory constraints. Various distributed fine-tuning methods have been proposed to alleviate memory constraints on GPU. However, determining the most effective method for achieving rapid fine-tuning while preventing GPU out-of-memory issues in a given environment remains unclear. To address this challenge, we introduce LLMem, a solution that estimates the GPU memory consumption when applying distributed fine-tuning methods across multiple GPUs and identifies the optimal method. We conduct GPU memory usage estimation prior to fine-tuning, leveraging the fundamental structure of transformer-based decoder models and the memory usage distribution of each method. Experimental results show that LLMem accurately estimates peak GPU memory usage on a single GPU, with error rates of up to 1.6%. Additionally, it shows an average error rate of 3.0% when applying distributed fine-tuning methods to LLMs with more than a billion parameters on multi-GPU setups.

cs.AI

MRGazer: Decoding Eye Gaze Points from Functional Magnetic Resonance Imaging in Individual Space

Eye-tracking research has proven valuable in understanding numerous cognitive functions. Recently, Frey et al. provided an exciting deep learning method for learning eye movements from fMRI data. However, it needed to co-register fMRI into standard space to obtain eyeballs masks, and thus required additional templates and was time consuming. To resolve this issue, in this paper, we propose a framework named MRGazer for predicting eye gaze points from fMRI in individual space. The MRGazer consisted of eyeballs extraction module and a residual network-based eye gaze prediction. Compared to the previous method, the proposed framework skips the fMRI co-registration step, simplifies the processing protocol and achieves end-to-end eye gaze regression. The proposed method achieved superior performance in a variety of eye movement tasks than the co-registration-based method, and delivered objective results within a shorter time (~ 0.02 Seconds for each volume) than prior method (~0.3 Seconds for each volume).

cs.CV

Multi-objective Generative Design of Three-Dimensional Composite Materials

Composite materials with 3D architectures are desirable in a variety of applications for the capability of tailoring their properties to meet multiple functional requirements. By the arrangement of materials' internal components, structure design is of great significance in tuning the properties of the composites. However, most of the composite structures are proposed by empirical designs following existing patterns. Hindered by the complexity of 3D structures, it is hard to extract customized structures with multiple desired properties from large design space. Here we report a multi-objective driven Wasserstein generative adversarial network (MDWGAN) to implement inverse designs of 3D composite structures according to given geometrical, structural and mechanical requirements. Our framework consists a GAN based network which generates 3D composite structures possessing with similar geometrical and structural features to the target dataset. Besides, multiple objectives are introduced to our framework for the control of mechanical property and isotropy of the composites. Real time calculation of the properties in training iterations is achieved by an accurate surrogate model. We constructed a small and concise dataset to illustrate our framework. With multiple objectives combined by their weight, and the 3D-GAN act as a soft constraint, our framework is proved to be capable of tuning the properties of the generated composites in multiple aspects, while keeping the selected features of different kinds of structures. The feasibility on small dataset and potential scalability on objectives of other properties make our work a novel, effective approach to provide fast, experience free composite structure designs for various functional materials.

cond-mat.mtrl-sci

The elastic properties of composites reinforced by a 3D Voronoi fibre network with or without missing fibres

Many composite materials, both natural and fabricated, process a Voronoi like architecture or microstructure. Furthermore, the stochasticity and connectivity of Voronoi tessellation endow the composite materials constructed by this kind of structures a wide range of desired properties including high and tuneable stiffness, high strength and good manufacturability. Thus, Voronoi-based fibre network structures are regarded as promising designs of composite reinforcements, as well as powerful tools in composite mechanics simulations. In this paper, the elastic properties of composites reinforced by a 3D Voronoi fibre network are systemically investigated based on the precise control of the Voronoi cell regularity. The regularity of the reinforcement fibre networks, according to our definition, is found positively related to the Young's moduli of the composites. Interestingly, 20% percent of defects in total reinforcement fibres only causes a less than 6.5% of Young's moduli drop in Voronoi fibre network reinforced composite. The Voronoi fibre network reinforced composites also shows higher Young's moduli than those of conventional composites with discrete reinforcements. The possibility of simulating aerogels by Voronoi fibre network structures are also presented.

cond-mat.mtrl-sci

Conformation-Induced Stiffening Effect of Crosslinked Polymer Thin Films

Nanoscale polymeric thin films are widely used in diverse applications such as energy devices, flexible electronics and biosensors, where a satisfactory mechanical performance is of vital importance to realize their full functionality. It has been evidenced that the elastic properties of polymer films are often strongly affected by their thickness; however, the underlying mechanism of this phenomenon, especially a thorough understanding at the microscopic level, has yet to be achieved. Here we established a coarse-grained molecular dynamics (CGMD) based computational framework, combining with experimental verifications, aiming to reveal the conformational origin of the stiffening behavior of crosslinked polymeric thin films. By imposing systematic controls over the polymer network structures, we found that the bi-axial modulus changes are essentially consequent of the alteration of polymer conformations. A unified theory was then proposed, to quantitatively clarify the correlation between the elastic properties of the system and the distributional variations of the chain end-to-end distances, with predicting a significant hardening effect on top of the conventional entropic elasticity with largely uncoiled chains. Adopting processing protocols inspired by the modeling, our experiments showed that PDMS films at approximately the same thickness may exhibit a two order of magnitude difference in their moduli. The good agreement between experiments and simulations illustrated our findings as an effective guideline for tailoring the elastic properties of polymer films at nanoscale.

cond-mat.soft

Revealing atomistic mechanisms of gold-catalyzed germanium growth using molecular dynamics simulations

The vapor-liquid-solid (VLS) method is considered a plausible technique for synthesizing germanium (Ge) nanostructures (e.g. nanowires), which have a broad range of applications due to their unique electronic properties and intrinsic compatibility with silicon. However, crystallization failures and material defects are still frequently observed in VLS processes, with insufficient understanding of their underlying mechanisms due to instrumental limitations for high-resolution in-situ characterizations. Employing an accurate interatomic potential well fitted to the gold-germanium (Au-Ge) phase diagram, we performed molecular dynamics simulations for a systematic investigation on the Au-catalyzed growth process of Ge crystals. From the simulations, relationships were established between the overall Ge growth rate and several main synthesis conditions, including substrate crystallographic orientation, temperature and Ge supersaturation in liquid. The dynamical behaviors of Ge atoms near the liquid-solid growing interface were captured, from which the atom surface stability and exchange rate were estimated for quantifying the atomistic details of the growth. These interface properties were further linked to the surface morphologies, to explain the observed orientation-dependent growing modes. This study sheds new lights into the understanding of the VLS growth mechanisms of Ge crystals, and provides scientific guidelines for designing innovative synthesis methods for similar nanomaterials.

cond-mat.mtrl-sci

Attention module improves both performance and interpretability of 4D fMRI decoding neural network

Decoding brain cognitive states from neuroimaging signals is an important topic in neuroscience. In recent years, deep neural networks (DNNs) have been recruited for multiple brain state decoding and achieved good performance. However, the open question of how to interpret the DNN black box remains unanswered. Capitalizing on advances in machine learning, we integrated attention modules into brain decoders to facilitate an in-depth interpretation of DNN channels. A 4D convolution operation was also included to extract temporo-spatial interaction within the fMRI signal. The experiments showed that the proposed model obtains a very high accuracy (97.4%) and outperforms previous researches on the 7 different task benchmarks from the Human Connectome Project (HCP) dataset. The visualization analysis further illustrated the hierarchical emergence of task-specific masks with depth. Finally, the model was retrained to regress individual traits within the HCP and to classify viewing images from the BOLD5000 dataset, respectively. Transfer learning also achieves good performance. A further visualization analysis shows that, after transfer learning, low-level attention masks remained similar to the source domain, whereas high-level attention masks changed adaptively. In conclusion, the proposed 4D model with attention module performed well and facilitated interpretation of DNNs, which is helpful for subsequent research.

eess.IV

Tuna: A Static Analysis Approach to Optimizing Deep Neural Networks

We introduce Tuna, a static analysis approach to optimizing deep neural network programs. The optimization of tensor operations such as convolutions and matrix multiplications is the key to improving the performance of deep neural networks. Many deep learning model optimization mechanisms today use dynamic analysis, which relies on experimental execution on a target device to build a data-driven cost model of the program. The reliance on dynamic profiling not only requires access to target hardware at compilation time but also incurs significant cost in machine resources. We introduce an approach that profiles the program by constructing features based on the target hardware characteristics in order. We use static analysis of the relative performance of tensor operations to optimize the deep learning program. Experiments show that our approach can achieve up to 11x performance compared to dynamic profiling based methods with the same compilation time.

cs.DC

Accelerating amorphous polymer electrolyte screening by learning to reduce errors in molecular dynamics simulated properties

Polymer electrolytes are promising candidates for the next generation lithium-ion battery technology. Large scale screening of polymer electrolytes is hindered by the significant cost of molecular dynamics (MD) simulation in amorphous systems: the amorphous structure of polymers requires multiple, repeated sampling to reduce noise and the slow relaxation requires long simulation time for convergence. Here, we accelerate the screening with a multi-task graph neural network that learns from a large amount of noisy, unconverged, short MD data and a small number of converged, long MD data. We achieve accurate predictions of 4 different converged properties and screen a space of 6247 polymers that is orders of magnitude larger than previous computational studies. Further, we extract several design principles for polymer electrolytes and provide an open dataset for the community. Our approach could be applicable to a broad class of material discovery problems that involve the simulation of complex, amorphous materials.

cond-mat.mtrl-sci