SearcharxivSearch

arXiv subjects

Lin Shen

Publications and source records attributed to Lin Shen.

6 recordsLinked to original sources

Fewest-Switches Surface Hopping with Combined Deep Learning Potential and Long Short-Term Memory Network Propagator for Simulating Realistic Photochemical Processes

Fewest-switches surface hopping (FSSH) is the most popular method for simulating photochemical processes of molecular systems. Recently, we have constructed long short-term memory (LSTM) networks as a propagator for electronic subsystems in FSSH dynamics simulations. The collective results on Tully's three models have been reproduced satisfactorily. In the present work, we develop an extended LSTM-FSSH framework to simulate realistic photochemical reactions. The input features of LSTM as well as the training procedure are redesigned to represent high-dimensional nuclear degrees of freedom in an effective way. Equivariant neural networks are integrated with LSTM to build adiabatic potential energy surfaces in ground and excited states. Photoisomerizations of $\mathrm{CH_2NH}$ and azobenzene are simulated, showing that our new proposed LSTM-FSSH method can produce excited-state lifetimes and product yields accurately in comparison with conventional FSSH simulations as reference. Only 10 reference trajectories are required for training LSTM networks, and then a trajectory ensemble can be generated with very efficient LSTM-FSSH dynamics simulations to obtain collective results.

physics.chem-ph

Uncovering and Aligning Anomalous Attention Heads to Defend Against NLP Backdoor Attacks

Backdoor attacks pose a serious threat to the security of large language models (LLMs), causing them to exhibit anomalous behavior under specific trigger conditions. The design of backdoor triggers has evolved from fixed triggers to dynamic or implicit triggers. This increased flexibility in trigger design makes it challenging for defenders to identify their specific forms accurately. Most existing backdoor defense methods are limited to specific types of triggers or rely on an additional clean model for support. To address this issue, we propose a backdoor detection method based on attention similarity, enabling backdoor detection without prior knowledge of the trigger. Our study reveals that models subjected to backdoor attacks exhibit unusually high similarity among attention heads when exposed to triggers. Based on this observation, we propose an attention safety alignment approach combined with head-wise fine-tuning to rectify potentially contaminated attention heads, thereby effectively mitigating the impact of backdoor attacks. Extensive experimental results demonstrate that our method significantly reduces the success rate of backdoor attacks while preserving the model's performance on downstream tasks.

cs.CR

EIRES:Training-free AI-Generated Image Detection via Edit-Induced Reconstruction Error Shift

Diffusion models have recently achieved remarkable photorealism, making it increasingly difficult to distinguish real images from generated ones, raising significant privacy and security concerns. In response, we present a key finding: structural edits enhance the reconstruction of real images while degrading that of generated images, creating a distinctive edit-induced reconstruction error shift. This asymmetric shift enhances the separability between real and generated images. Building on this insight, we propose EIRES, a training-free method that leverages structural edits to reveal inherent differences between real and generated images. To explain the discriminative power of this shift, we derive the reconstruction error lower bound under edit perturbations. Since EIRES requires no training, thresholding depends solely on the natural separability of the signal, where a larger margin yields more reliable detection. Extensive experiments show that EIRES is effective across diverse generative models and remains robust on the unbiased subset, even under post-processing operations.

cs.CV

Large Language Models Illuminate a Progressive Pathway to Artificial Healthcare Assistant: A Review

With the rapid development of artificial intelligence, large language models (LLMs) have shown promising capabilities in mimicking human-level language comprehension and reasoning. This has sparked significant interest in applying LLMs to enhance various aspects of healthcare, ranging from medical education to clinical decision support. However, medicine involves multifaceted data modalities and nuanced reasoning skills, presenting challenges for integrating LLMs. This paper provides a comprehensive review on the applications and implications of LLMs in medicine. It begins by examining the fundamental applications of general-purpose and specialized LLMs, demonstrating their utilities in knowledge retrieval, research support, clinical workflow automation, and diagnostic assistance. Recognizing the inherent multimodality of medicine, the review then focuses on multimodal LLMs, investigating their ability to process diverse data types like medical imaging and EHRs to augment diagnostic accuracy. To address LLMs' limitations regarding personalization and complex clinical reasoning, the paper explores the emerging development of LLM-powered autonomous agents for healthcare. Furthermore, it summarizes the evaluation methodologies for assessing LLMs' reliability and safety in medical contexts. Overall, this review offers an extensive analysis on the transformative potential of LLMs in modern medicine. It also highlights the pivotal need for continuous optimizations and ethical oversight before these models can be effectively integrated into clinical practice. Visit https://github.com/mingze-yuan/Awesome-LLM-Healthcare for an accompanying GitHub repository containing latest papers.

cs.CL

Fewest-Switches Surface Hopping with Long Short-Term Memory Networks

The mixed quantum-classical dynamical simulation is essential to study nonadiabatic phenomena in photophysics and photochemistry. In recent years, many machine learning models have been developed to accelerate the time evolution of the nuclear subsystem. Herein, we implement long short-term memory (LSTM) networks as a propagator to accelerate the time evolution of the electronic subsystem during the fewest-switches surface hopping (FSSH) simulations. A small number of reference trajectories are generated using the original FSSH method, and then the LSTM networks can be built, accompanied by careful examination of typical LSTM-FSSH trajectories that employ the same initial condition and random numbers as the corresponding reference. The constructed network is applied to FSSH to further produce a trajectory ensemble to reveal the mechanism of nonadiabatic processes. Taking Tully's three models as test systems, the collective results can be reproduced qualitatively. This work demonstrates that LSTM is applicable to the most popular surface hopping simulations.

physics.chem-ph

Application of machine learning techniques at BESIII experiment

BESIII is a currently running tau-charm factory with the largest samples of on threshold charm meson pairs, directly produced charmonia and some other unique datasets at BEPCII collider. Machine learning techniques have been employed to improve the performance of BESIII software. The studies for reweighing MC, particle identification and cluster reconstruction for the CGEM (Cylindrical Gas Electron Multiplier) inner tracker are presented.

physics.ins-det