SearcharxivSearch

arXiv subjects

Xiangrui Gao

Publications and source records attributed to Xiangrui Gao.

5 recordsLinked to original sources

siDPT: siRNA Efficacy Prediction via Debiased Preference-Pair Transformer

Small interfering RNA (siRNA) is a short double-stranded RNA molecule (about 21-23 nucleotides) with the potential to cure diseases by silencing the function of target genes. Due to its well-understood mechanism, many siRNA-based drugs have been evaluated in clinical trials. However, selecting effective binding regions and designing siRNA sequences requires extensive experimentation, making the process costly. As genomic resources and publicly available siRNA datasets continue to grow, data-driven models can be leveraged to better understand siRNA-mRNA interactions. To fully exploit such data, curating high-quality siRNA datasets is essential to minimize experimental errors and noise. We propose siDPT: siRNA efficacy Prediction via Debiased Preference-Pair Transformer, a framework that constructs a preference-pair dataset and designs an siRNA-mRNA interactive transformer with debiased ranking objectives to improve siRNA inhibition prediction and generalization. We evaluate our approach using two public datasets and one newly collected patent dataset. Our model demonstrates substantial improvement in Pearson correlation and strong performance across other metrics.

q-bio.GN

mRNA2vec: mRNA Embedding with Language Model in the 5'UTR-CDS for mRNA Design

Messenger RNA (mRNA)-based vaccines are accelerating the discovery of new drugs and revolutionizing the pharmaceutical industry. However, selecting particular mRNA sequences for vaccines and therapeutics from extensive mRNA libraries is costly. Effective mRNA therapeutics require carefully designed sequences with optimized expression levels and stability. This paper proposes a novel contextual language model (LM)-based embedding method: mRNA2vec. In contrast to existing mRNA embedding approaches, our method is based on the self-supervised teacher-student learning framework of data2vec. We jointly use the 5' untranslated region (UTR) and coding sequence (CDS) region as the input sequences. We adapt our LM-based approach specifically to mRNA by 1) considering the importance of location on the mRNA sequence with probabilistic masking, 2) using Minimum Free Energy (MFE) prediction and Secondary Structure (SS) classification as additional pretext tasks. mRNA2vec demonstrates significant improvements in translation efficiency (TE) and expression level (EL) prediction tasks in UTR compared to SOTA methods such as UTR-LM. It also gives a competitive performance in mRNA stability and protein production level tasks in CDS such as CodonBERT.

q-bio.QM

StarGraph: Knowledge Representation Learning based on Incomplete Two-hop Subgraph

Conventional representation learning algorithms for knowledge graphs (KG) map each entity to a unique embedding vector, ignoring the rich information contained in the neighborhood. We propose a method named StarGraph, which gives a novel way to utilize the neighborhood information for large-scale knowledge graphs to obtain entity representations. An incomplete two-hop neighborhood subgraph for each target node is at first generated, then processed by a modified self-attention network to obtain the entity representation, which is used to replace the entity embedding in conventional methods. We achieved SOTA performance on ogbl-wikikg2 and got competitive results on fb15k-237. The experimental results proves that StarGraph is efficient in parameters, and the improvement made on ogbl-wikikg2 demonstrates its great effectiveness of representation learning on large-scale knowledge graphs. The code is now available at \url{https://github.com/hzli-ucas/StarGraph}.

cs.CL

Labelled tree graphs, Feynman diagrams and disk integrals

In this note, we introduce and study a new class of "half integrands" in Cachazo-He-Yuan (CHY) formula, which naturally generalize the so-called Parke-Taylor factors; these are dubbed Cayley functions as each of them corresponds to a labelled tree graph. The CHY formula with a Cayley function squared gives a sum of Feynman diagrams, and we represent it by a combinatoric polytope whose vertices correspond to Feynman diagrams. We provide a simple graphic rule to derive the polytope from a labelled tree graph, and classify such polytopes ranging from the associahedron to the permutohedron. Furthermore, we study the linear space of such half integrands and find (1) a nice formula reducing any Cayley function to a sum of Parke-Taylor factors in the Kleiss-Kuijf basis (2) a set of Cayley functions as a new basis of the space; each element has the remarkable property that its CHY formula with a given Parke-Taylor factor gives either a single Feynman diagram or zero. We also briefly discuss applications of Cayley functions and the new basis in certain disk integrals of superstring theory.

hep-th

Relativistic correction to gluon fragmentation function into pseudoscalar quarkonium

Inspired by the recent measurements of the $η_c$ meson production at LHC, we investigate the relativistic correction effect for the fragmentation function of the gluon into $η_c$, which constitutes the crucial nonperturbative elements to understand $η_c$ production at high $p_T$. Employing three distinct methods, we calculate the leading relativistic correction to the $g\toη_c$ fragmentation function in the NRQCD factorization framework, as well as verify the existing NLO result for the $c\to η_c$ fragmentation function. We also study the evolution behavior of these fragmentation functions with the aid of DGLAP equation.

hep-ph