SearcharxivSearch

arXiv subjects

Jike Wang

Publications and source records attributed to Jike Wang.

14 recordsLinked to original sources

Low-energy Muon-Nucleon scattering experiment: LUNE (White Paper)

The HIAF will provide high-intensity, high-quality muon beams with momenta from 0.5 to 7.5 GeV/c. This energy range is uniquely suited for precision muon scattering, bridging the gap between low-energy electron facilities and future high-energy lepton-ion colliders. In particular, HIAF will enable precision measurements with both positive and negative muon beams over a broad kinematic range, complementing existing electron-scattering facilities such as JLab, EicC and EIC. Based on HIAF muon source, the LUNE Collaboration has been established to address several fundamental questions in nuclear and particle physics, including the proton charge radius puzzle, nucleon electromagnetic structure, and the dynamics of quantum electrodynamics and hadronic interactions. The program proceeds in two phases, from elastic scattering to nucleon structure and beyond-Standard-Model searches. The experiment is expected to determine the proton charge radius with a precision of approximately 1.0\% using elastic muon-proton scattering. It will also perform systematic measurements of the proton electromagnetic form factors with both $\mu^+$ and $\mu^-$ beams, enabling precise studies of two-photon exchange effects and stringent tests of quantum electrodynamics. Beyond elastic scattering, LUNE will investigate TMD, gravitational form factors, and nuclear charge radii, providing new insights into the 3D structure of nucleons and nuclei. The experiment will further address important topics including Coulomb-distortion corrections, nuclear medium effects, and possible signatures of physics beyond the Standard Model. This white paper presents the scientific motivation, detector concept, expected performance, and long-term strategy of LUNE.

hep-ex

Rethinking Diffusion Models with Symmetries through Canonicalization with Applications to Molecular Graph Generation

Many generative tasks in chemistry and science involve distributions invariant to group symmetries (e.g., permutation and rotation). A common strategy enforces invariance and equivariance through architectural constraints such as equivariant denoisers and invariant priors. In this paper, we challenge this tradition through the alternative canonicalization perspective: first map each sample to an orbit representative with a canonical pose or order, train an unconstrained (non-equivariant) diffusion or flow model on the canonical slice, and finally recover the invariant distribution by sampling a random symmetry transform at generation time. Building on a formal quotient-space perspective, our work provides a comprehensive theory of canonical diffusion by proving: (i) the correctness, universality and superior expressivity of canonical generative models over invariant targets; (ii) canonicalization accelerates training by removing diffusion score complexity induced by group mixtures and reducing conditional variance in flow matching. We then show that aligned priors and optimal transport act complementarily with canonicalization and further improves training efficiency. We instantiate the framework for molecular graph generation under $S_n \times SE(3)$ symmetries. By leveraging geometric spectra-based canonicalization and mild positional encodings, canonical diffusion significantly outperforms equivariant baselines in 3D molecule generation tasks, with similar or even less computation. Moreover, with a novel architecture Canon, CanonFlow achieves state-of-the-art performance on the challenging GEOM-DRUG dataset, and the advantage remains large in few-step generation.

cs.LG

{\mu}Touch: Enabling Accurate, Lightweight Self-Touch Sensing with Passive Magnets

Self-touch gestures (e.g., nuanced facial touches and subtle finger scratches) provide rich insights into human behaviors, from hygiene practices to health monitoring. However, existing approaches fall short in detecting such micro gestures due to their diverse movement patterns. This paper presents {\mu}Touch, a novel magnetic sensing platform for self-touch gesture recognition. {\mu}Touch features (1) a compact hardware design with low-power magnetometers and magnetic silicon, (2) a lightweight semi-supervised framework requiring minimal user data, and (3) an ambient field detection module to mitigate environmental interference. We evaluated {\mu}Touch in two representative applications in user studies with 11 and 12 participants. {\mu}Touch only requires three-second fine-tuning data for each gesture, and new users need less than one minute before starting to use the system. {\mu}Touch can distinguish eight different face-touching behaviors with an average accuracy of 93.41%, and reliably detect body-scratch behaviors with an average accuracy of 94.63%. {\mu}Touch demonstrates accurate and robust sensing performance even after a month, showcasing its potential as a practical tool for hygiene monitoring and dermatological health applications. Code is available at https://wangmerlyn.github.io/muTouch/.

cs.HC

A Scalable and Quantum-Accurate Foundation Model for Biomolecular Force Field via Linearly Tensorized Quadrangle Attention

Accurate atomistic biomolecular simulations are vital for disease mechanism understanding, drug discovery, and biomaterial design, but existing simulation methods exhibit significant limitations. Classical force fields are efficient but lack accuracy for transition states and fine conformational details critical in many chemical and biological processes. Quantum Mechanics (QM) methods are highly accurate but computationally infeasible for large-scale or long-time simulations. AI-based force fields (AIFFs) aim to achieve QM-level accuracy with efficiency but struggle to balance many-body modeling complexity, accuracy, and speed, often constrained by limited training data and insufficient validation for generalizability. To overcome these challenges, we introduce LiTEN, a novel equivariant neural network with Tensorized Quadrangle Attention (TQA). TQA efficiently models three- and four-body interactions with linear complexity by reparameterizing high-order tensor features via vector operations, avoiding costly spherical harmonics. Building on LiTEN, LiTEN-FF is a robust AIFF foundation model, pre-trained on the extensive nablaDFT dataset for broad chemical generalization and fine-tuned on SPICE for accurate solvated system simulations. LiTEN achieves state-of-the-art (SOTA) performance across most evaluation subsets of rMD17, MD22, and Chignolin, outperforming leading models such as MACE, NequIP, and EquiFormer. LiTEN-FF enables the most comprehensive suite of downstream biomolecular modeling tasks to date, including QM-level conformer searches, geometry optimization, and free energy surface construction, while offering 10x faster inference than MACE-OFF for large biomolecules (~1000 atoms). In summary, we present a physically grounded, highly efficient framework that advances complex biomolecular modeling, providing a versatile foundation for drug discovery and related applications.

physics.chem-ph

Graph Neural Networks in Modern AI-aided Drug Discovery

Graph neural networks (GNNs), as topology/structure-aware models within deep learning, have emerged as powerful tools for AI-aided drug discovery (AIDD). By directly operating on molecular graphs, GNNs offer an intuitive and expressive framework for learning the complex topological and geometric features of drug-like molecules, cementing their role in modern molecular modeling. This review provides a comprehensive overview of the methodological foundations and representative applications of GNNs in drug discovery, spanning tasks such as molecular property prediction, virtual screening, molecular generation, biomedical knowledge graph construction, and synthesis planning. Particular attention is given to recent methodological advances, including geometric GNNs, interpretable models, uncertainty quantification, scalable graph architectures, and graph generative frameworks. We also discuss how these models integrate with modern deep learning approaches, such as self-supervised learning, multi-task learning, meta-learning and pre-training. Throughout this review, we highlight the practical challenges and methodological bottlenecks encountered when applying GNNs to real-world drug discovery pipelines, and conclude with a discussion on future directions.

q-bio.BM

AutoLoop: a novel autoregressive deep learning method for protein loop prediction with high accuracy

Protein structure prediction is a critical and longstanding challenge in biology, garnering widespread interest due to its significance in understanding biological processes. A particular area of focus is the prediction of missing loops in proteins, which are vital in determining protein function and activity. To address this challenge, we propose AutoLoop, a novel computational model designed to automatically generate accurate loop backbone conformations that closely resemble their natural structures. AutoLoop employs a bidirectional training approach while merging atom- and residue-level embedding, thus improving robustness and precision. We compared AutoLoop with twelve established methods, including FREAD, NGK, AlphaFold2, and AlphaFold3. AutoLoop consistently outperforms other methods, achieving a median RMSD of 1.12 Angstrom and a 2-Angstrom success rate of 73.23% on the CASP15 dataset, while maintaining strong performance on the HOMSTARD dataset. It demonstrates the best performance across nearly all loop lengths and secondary structural types. Beyond accuracy, AutoLoop is computationally efficient, requiring only 0.10 s per generation. A post-processing module for side-chain packing and energy minimization further improves results slightly, confirming the reliability of the predicted backbone. A case study also highlights AutoLoop's potential for precise predictions based on dominant loop conformations. These advances hold promise for protein engineering and drug discovery.

q-bio.BM

Advancing biomolecular understanding and design following human instructions

Understanding and designing biomolecules, such as proteins and small molecules, is central to advancing drug discovery, synthetic biology and enzyme engineering. Recent breakthroughs in artificial intelligence have revolutionized biomolecular research, achieving remarkable accuracy in biomolecular prediction and design. However, a critical gap remains between artificial intelligence's computational capabilities and researchers' intuitive goals, particularly in using natural language to bridge complex tasks with human intentions. Large language models have shown potential to interpret human intentions, yet their application to biomolecular research remains nascent due to challenges including specialized knowledge requirements, multimodal data integration, and semantic alignment between natural language and biomolecules. To address these limitations, we present InstructBioMol, a large language model designed to bridge natural language and biomolecules through a comprehensive any-to-any alignment of natural language, molecules and proteins. This model can integrate multimodal biomolecules as the input, and enable researchers to articulate design goals in natural language, providing biomolecular outputs that meet precise biological needs. Experimental results demonstrate that InstructBioMol can understand and design biomolecules following human instructions. In particular, it can generate drug molecules with a 10% improvement in binding affinity and design enzymes that achieve an enzyme-substrate pair prediction score of 70.4. This highlights its potential to transform real-world biomolecular research. The code is available at https://github.com/HICAI-ZJU/InstructBioMol.

cs.CL

Token-Mol 1.0: Tokenized drug design with large language model

Significant interests have recently risen in leveraging sequence-based large language models (LLMs) for drug design. However, most current applications of LLMs in drug discovery lack the ability to comprehend three-dimensional (3D) structures, thereby limiting their effectiveness in tasks that explicitly involve molecular conformations. In this study, we introduced Token-Mol, a token-only 3D drug design model. This model encodes all molecular information, including 2D and 3D structures, as well as molecular property data, into tokens, which transforms classification and regression tasks in drug discovery into probabilistic prediction problems, thereby enabling learning through a unified paradigm. Token-Mol is built on the transformer decoder architecture and trained using random causal masking techniques. Additionally, we proposed the Gaussian cross-entropy (GCE) loss function to overcome the challenges in regression tasks, significantly enhancing the capacity of LLMs to learn continuous numerical values. Through a combination of fine-tuning and reinforcement learning (RL), Token-Mol achieves performance comparable to or surpassing existing task-specific methods across various downstream tasks, including pocket-based molecular generation, conformation generation, and molecular property prediction. Compared to existing molecular pre-trained models, Token-Mol exhibits superior proficiency in handling a wider range of downstream tasks essential for drug design. Notably, our approach improves regression task accuracy by approximately 30% compared to similar token-only methods. Token-Mol overcomes the precision limitations of token-only models and has the potential to integrate seamlessly with general models such as ChatGPT, paving the way for the development of a universal artificial intelligence drug design model that facilitates rapid and high-quality drug design by experts.

q-bio.BM

Prediction of the treatment effect of FLASH radiotherapy with Circular Electron-Positron Collider (CEPC) synchrotron radiation

The Circular Electron-Positron Collider (CEPC) can also work as a powerful and excellent synchrotron light source, which can generate high-quality synchrotron radiation. This synchrotron radiation has potential advantages in the medical field, with a broad spectrum, with energies ranging from visible light to x-rays used in conventional radiotherapy, up to several MeV. FLASH radiotherapy is one of the most advanced radiotherapy modalities. It is a radiotherapy method that uses ultra-high dose rate irradiation to achieve the treatment dose in an instant; the ultra-high dose rate used is generally greater than 40 Gy/s, and this type of radiotherapy can protect normal tissues well. In this paper, the treatment effect of CEPC synchrotron radiation for FLASH radiotherapy was evaluated by simulation. First, Geant4 simulation was used to build a synchrotron radiation radiotherapy beamline station, and then the dose rate that CEPC can produce was calculated. Then, a physicochemical model of radiotherapy response kinetics was established, and a large number of radiotherapy experimental data were comprehensively used to fit and determine the functional relationship between the treatment effect, dose rate and dose. Finally, the macroscopic treatment effect of FLASH radiotherapy was predicted using CEPC synchrotron radiation light through the dose rate and the above-mentioned functional relationship. The results show that CEPC synchrotron radiation beam is one of the best beams for FLASH radiotherapy.

physics.med-ph

Discovery of novel antimicrobial peptides with notable antibacterial potency by a LLM-based foundation model

Large language models (LLMs) have shown remarkable advancements in chemistry and biomedical research, acting as versatile foundation models for various tasks. We introduce AMP-Designer, an LLM-based approach for swiftly designing novel antimicrobial peptides (AMPs) with desired properties. Within 11 days, AMP-Designer achieved the de novo design of 18 AMPs with broad-spectrum activity against Gram-negative bacteria. In vitro validation revealed a 94.4% success rate, with two candidates demonstrating exceptional antibacterial efficacy, minimal hemotoxicity, stability in human plasma, and low potential to induce resistance, as evidenced by significant bacterial load reduction in murine lung infection experiments. The entire process, from design to validation, concluded in 48 days. AMP-Designer excels in creating AMPs targeting specific strains despite limited data availability, with a top candidate displaying a minimum inhibitory concentration of 2.0 {\mu}g/ml against Propionibacterium acnes. Integrating advanced machine learning techniques, AMP-Designer demonstrates remarkable efficiency, paving the way for innovative solutions to antibiotic resistance.

q-bio.BM

Angular asymmetries in $B\toΛ\bar p M$ decays

The forward-backward angular asymmetry (${\cal A}_{FB}$) for $\bar B^0\to Λ\bar pπ^+$ measured by Belle has presented an experimental value in the range of $-30\%$ to $-50\%$. In our study, we find that ${\cal A}_{FB}[\bar B^0\toΛ\bar p π^+(B^-\to Λ\bar pπ^0)]$ can be as large as $(-14.6^{+0.9}_{-1.5}\pm 6.9)\%$. In addition, we present ${\cal A}_{FB}[\bar B^0\toΛ\bar p ρ^+(B^-\to Λ\bar pρ^0)] =(4.1^{+2.8}_{-0.7}\pm 2.0)\%$ as the first prediction involving a vector meson in the charmless $B\to{\bf B\bar B'}M$ decays. While ${\cal A}_{FB}(B\toΛ\bar p M)$ indicates an angular correlation caused by the rarely studied baryonic form factors in the timelike region, LHCb and Belle~II are capable of performing experimental examinations.

hep-ph

Baryonic $B$ meson decays

We review the two and three-body baryonic $B$ decays with the dibaryon (${\bf B\bar B'}$) as the final states. Accordingly, we summarize the experimental data of the branching fractions, angular asymmetries, and $CP$ asymmetries. Using the $W$-boson annihilation (exchange) mechanism, the branching fractions of $B\to {\bf B \bf \bar B'}$ are shown to be interpretable. In the approach of perturbative QCD counting rules, we study the three-body decay channels. In particular, we review the $CP$ asymmetries of $B\to {\bf B\bar B'}M$, which are promising to be measured by the LHCb and Belle~II experiments. Finally, we remark the theoretical challenges in interpreting ${\cal B}(B^-\to p\bar pρ^-)$ and ${\cal B}(B^-\to p\bar pμ^-\bar ν_μ)$.

hep-ph

Triplet Track Trigger for Future High Rate Experiments

The hadron-hadron based Future Circular Collider (FCC-hh) is a project with the goal to collide proton-proton beams at a center of mass energy of sqrt(s)=100 TeV with a bunch crossing rate of 25 ns. Some of the major challenges that the FCC-experiments have to tackle are the very large number of pile-up events of around 1000 and the data processing, namely the reduction of the huge data rate of 1-2 PBytes/s whilst keeping the signal efficiencies of interesting processes high. Therefore, smart triggering concepts are needed that not only allow for a significant reduction of pile-up and rate but also provide high signal acceptance and purity. In this proceeding, one such concept of a triplet track trigger based on Monolithic ActivePixel Sensors (MAPS) is presented for a generic detector geometry. It is demonstrated that the triplet pixel layer design allows for a very simple and fast track reconstruction already at the first trigger level, providing excellent track reconstruction efficiencies and very high purity at the same time. Based on a full Geant4 simulation tracking performance studies are presented for a full-scale triplet pixel detector, i.e. three closely spaced pixel layers at a sufficiently large radius, in an FCC-like detector environment. Results obtained for different triplet layer design parameters are compared.

physics.ins-det

Simulation for the ATLAS Upgrade Strip Tracker

ATLAS is making extensive efforts towards preparing a detector upgrade for the high luminosity operations of the LHC (HL-LHC), which will commence operation in about 10 years. The current ATLAS Inner Detector will be replaced by an all-silicon tracker (comprising an inner Pixel tracker and outer Strip tracker). The software currently used for the new silicon tracker is broadly inherited from that used for the LHC Run-1 and Run-2, but many new developments have been made to better fulfill the future detector and operation requirements. One aspect in particular which will be highlighted is the simulation software for the Strip tracker. The available geometry description software (including the detailed description for all the sensitive elements, the services, etc.) did not allow for accurate modelling of the planned detector design. A range of sensors/layouts for the Strip tracker are being considered and must be studied in detailed simulations in order to assess the performance and ascertain that requirements are met. For this, highly flexibility geometry building is required from the simulation software. A new Xml-based detector description framework has been developed to meet the aforementioned challenges. We will present the design of the framework and its validation results.

physics.ins-det