SearcharxivSearch

arXiv subjects

Yeqing Lu

Publications and source records attributed to Yeqing Lu.

3 recordsLinked to original sources

PathoMIC: A Benchmark for Cross-Species Antimicrobial Peptide Activity Prediction

With activity against multidrug-resistant pathogens and mechanisms distinct from conventional antibiotics, antimicrobial peptides (AMPs) offer a promising approach to combating antibiotic-resistant infections. However, their potency varies substantially across pathogen species, making accurate prediction of the minimum inhibitory concentration (MIC) for specific peptide-pathogen pairs essential for prioritizing candidates before costly experimental validation. Existing predictors are trained mainly on a few well-represented pathogens and rarely exploit biological relationships across species, limiting their generalization to low-resource and unseen pathogens. We introduce PathoMIC, the largest and most pathogen-diverse unified dataset for quantitative antimicrobial peptide activity prediction, containing 74,751 experimentally reported MIC measurements across 424 pathogen species. PathoMIC integrates peptide sequences, standardized MIC values, pathogen descriptions, and taxonomic relationships to facilitate knowledge transfer across related species. We establish few-shot and zero-shot cross-species evaluation protocols and develop a knowledge-enhanced framework that leverages pathogen descriptions and taxonomy. The framework yields substantial improvements for low-resource species with limited supervision, while gains for entirely unseen species remain modest, highlighting the difficulty of zero-shot cross-species MIC prediction. PathoMIC provides a standardized foundation for cross-species activity modeling and pathogen-specific virtual screening. Code is available at https://anonymous.4open.science/r/PathoMIC-546D/.

q-bio.QM

Nature Language Model: Deciphering the Language of Nature for Scientific Discovery

Foundation models have revolutionized natural language processing and artificial intelligence, significantly enhancing how machines comprehend and generate human languages. Inspired by the success of these foundation models, researchers have developed foundation models for individual scientific domains, including small molecules, materials, proteins, DNA, RNA and even cells. However, these models are typically trained in isolation, lacking the ability to integrate across different scientific domains. Recognizing that entities within these domains can all be represented as sequences, which together form the "language of nature", we introduce Nature Language Model (NatureLM), a sequence-based science foundation model designed for scientific discovery. Pre-trained with data from multiple scientific domains, NatureLM offers a unified, versatile model that enables various applications including: (i) generating and optimizing small molecules, proteins, RNA, and materials using text instructions; (ii) cross-domain generation/design, such as protein-to-molecule and protein-to-RNA generation; and (iii) top performance across different domains, matching or surpassing state-of-the-art specialist models. NatureLM offers a promising generalist approach for various scientific tasks, including drug discovery (hit generation/optimization, ADMET optimization, synthesis), novel material design, and the development of therapeutic proteins or nucleotides. We have developed NatureLM models in different sizes (1 billion, 8 billion, and 46.7 billion parameters) and observed a clear improvement in performance as the model size increases.

cs.AI

Scalability of Atomic-Thin-Body (ATB) Transistors Based on Graphene Nanoribbons

A general solution for the electrostatic potential in an atomic-thin-body (ATB) field-effect transistor geometry is presented. The effective electrostatic scaling length, λeff, is extracted from the analytical model, which cannot be approximated by the lowest order eigenmode as traditionally done in SOI-MOSFETs. An empirical equation for the scaling length that depends on the geometry parameters is proposed. It is shown that even for a thick SiO2 back oxide λeff can be improved efficiently by thinner top oxide thickness, and to some extent, with high-k dielectrics. The model is then applied to self-consistent simulation of graphene nanoribbon (GNR) Schottky-barrier field-effect transistors (SB-FETs) at the ballistic limit. In the case of GNR SB-FETs, for large λeff, the scaling is limited by the conventional electrostatic short channel effects (SCEs). On the other hand, for small λeff, the scaling is limited by direct source-to-drain tunneling. A subthreshold swing below 100mV/dec is still possible with a sub-10nm gate length in GNR SB-FETs.

cond-mat.mes-hall