SearcharxivSearch

arXiv subjects

Mengjie Chen

Publications and source records attributed to Mengjie Chen.

12 recordsLinked to original sources

MGKAN: Predicting Asymmetric Drug-Drug Interactions via a Multimodal Graph Kolmogorov-Arnold Network

Predicting drug-drug interactions (DDIs) is essential for safe pharmacological treatments. Previous graph neural network (GNN) models leverage molecular structures and interaction networks but mostly rely on linear aggregation and symmetric assumptions, limiting their ability to capture nonlinear and heterogeneous patterns. We propose MGKAN, a Graph Kolmogorov-Arnold Network that introduces learnable basis functions into asymmetric DDI prediction. MGKAN replaces conventional MLP transformations with KAN-driven basis functions, enabling more expressive and nonlinear modeling of drug relationships. To capture pharmacological dependencies, MGKAN integrates three network views-an asymmetric DDI network, a co-interaction network, and a biochemical similarity network-with role-specific embeddings to preserve directional semantics. A fusion module combines linear attention and nonlinear transformation to enhance representational capacity. On two benchmark datasets, MGKAN outperforms seven state-of-the-art baselines. Ablation studies and case studies confirm its predictive accuracy and effectiveness in modeling directional drug effects.

cs.LG

Narrowline Laser Cooling and Spectroscopy of Molecules via Stark States

The electronic energy level structure of yttrium monoxide (YO) provides a long-lived, low-lying $^{2}Δ$ state ideal for high-precision molecular spectroscopy, narrowline laser cooling at the single photon-recoil limit, and studying dipolar physics with unprecedented interaction strength. High-resolution laser spectroscopy of ultracold laser-cooled YO molecules is used to study the Stark effect in the A$^{\prime}\,^{2}Δ_{3/2}\,J=3/2$ state. An immediate onset of the linear Stark effect is observed in the presence of weak applied electric fields due to the near degenerate $Λ$-doublet and the large electric dipole moment. By applying a small electric field the Stark insensitive state is spectroscopically isolated and the absolute transition frequency to the X$\,^2Σ^+$ electronic ground state is determined with a fractional frequency uncertainty of 9 $\times$ 10$^{-12}$. This electric field control is necessary to implement a quasi-closed photon cycling scheme that preserves parity. With this scheme the first narrowline laser cooling of a molecules is demonstrated, reducing the temperature of sub-Doppler cooled YO in two dimensions.

physics.atom-ph

Towards Interpretable Drug-Drug Interaction Prediction: A Graph-Based Approach with Molecular and Network-Level Explanations

Drug-drug interactions (DDIs) represent a critical challenge in pharmacology, often leading to adverse drug reactions with significant implications for patient safety and healthcare outcomes. While graph-based methods have achieved strong predictive performance, most approaches treat drug pairs independently, overlooking the complex, context-dependent interactions unique to drug pairs. Additionally, these models struggle to integrate biological interaction networks and molecular-level structures to provide meaningful mechanistic insights. In this study, we propose MolecBioNet, a novel graph-based framework that integrates molecular and biomedical knowledge for robust and interpretable DDI prediction. By modeling drug pairs as unified entities, MolecBioNet captures both macro-level biological interactions and micro-level molecular influences, offering a comprehensive perspective on DDIs. The framework extracts local subgraphs from biomedical knowledge graphs and constructs hierarchical interaction graphs from molecular representations, leveraging classical graph neural network methods to learn multi-scale representations of drug pairs. To enhance accuracy and interpretability, MolecBioNet introduces two domain-specific pooling strategies: context-aware subgraph pooling (CASPool), which emphasizes biologically relevant entities, and attention-guided influence pooling (AGIPool), which prioritizes influential molecular substructures. The framework further employs mutual information minimization regularization to enhance information diversity during embedding fusion. Experimental results demonstrate that MolecBioNet outperforms state-of-the-art methods in DDI prediction, while ablation studies and embedding visualizations further validate the advantages of unified drug pair modeling and multi-scale knowledge integration.

cs.LG

Collisions of Spin-polarized YO Molecules for Single Partial Waves

Efficient sub-Doppler laser cooling and optical trapping of YO molecules offer new opportunities to study collisional dynamics in the quantum regime. Confined in a crossed optical dipole trap, we achieve the highest phase-space density of $2.5 \times 10^{-5}$ for a bulk laser-cooled molecular sample. This sets the stage to study YO--YO collisions in the microkelvin temperature regime, and reveal state-dependent, single-partial-wave two-body collisional loss rates. We determine the partial-wave contributions to loss of specific rotational states (first excited $N=1$ and ground $N=0$) following two strategies. First, we measure the change of the collision rate in a spin mixture of $N=1$ by tuning the kinetic energy with respect to the p- and d-wave centrifugal barriers. Second, we compare loss rates between a spin mixture and a spin-polarized state in $N=0$. Using quantum defect theory with a partially absorbing boundary condition at short range, we show that the dependence on temperature for $N=1$ can be reproduced in the presence of a d-wave or f-wave resonance, and the dependence on a spin mixture for $N=0$ with a p-wave resonance.

physics.atom-ph

AGChain: A Blockchain-based Gateway for Trustworthy App Delegation from Mobile App Markets

The popularity of smartphones has led to the growth of mobile app markets, creating a need for enhanced transparency, global access, and secure downloading. This paper introduces AGChain, a blockchain-based gateway that enables trustworthy app delegation within existing markets. AGChain ensures that markets can continue providing services while users benefit from permanent, distributed, and secure app delegation. During its development, we address two key challenges: significantly reducing smart contract gas costs and enabling fully distributed IPFS-based file storage. Additionally, we tackle three system issues related to security and sustainability. We have implemented a prototype of AGChain on Ethereum and Polygon blockchains, achieving effective security and decentralization with a minimal gas cost of around 0.002 USD per app upload (no cost for app download). The system also exhibits reasonable performance with an average overhead of 12%.

cs.CR

The construction of ceRNAs network reveals the prognostic characteristics of prostate cancer

The dysregulation of transcripts is characterized as one of the main mechanisms in tumor pathogenesis. The recent discovery developed a new hypothesis, competitive endogenous RNAs (ceRNAs), which could regulate other RNA transcripts via competing for their shared miRNAs. The interaction of elements in ceRNAs network was involved in a large range of biological reactions and facilitate to cancer progression. In this study, we performed a comprehensive investigation on the regulatory mechanisms and functional roles of ceRNAs in prostate cancer (PCa) and constructed a ceRNAs network which could possess potential value in patient prognosis and be evaluated as therapeutic targets for PCa.

q-bio.MN

Robust Covariance and Scatter Matrix Estimation under Huber's Contamination Model

Covariance matrix estimation is one of the most important problems in statistics. To accommodate the complexity of modern datasets, it is desired to have estimation procedures that not only can incorporate the structural assumptions of covariance matrices, but are also robust to outliers from arbitrary sources. In this paper, we define a new concept called matrix depth and then propose a robust covariance matrix estimator by maximizing the empirical depth function. The proposed estimator is shown to achieve minimax optimal rate under Huber's $ε$-contamination model for estimating covariance/scatter matrices with various structures including bandedness and sparsity.

math.ST

A General Decision Theory for Huber's $ε$-Contamination Model

Today's data pose unprecedented challenges to statisticians. It may be incomplete, corrupted or exposed to some unknown source of contamination. We need new methods and theories to grapple with these challenges. Robust estimation is one of the revived fields with potential to accommodate such complexity and glean useful information from modern datasets. Following our recent work on high dimensional robust covariance matrix estimation, we establish a general decision theory for robust statistics under Huber's $ε$-contamination model. We propose a solution using Scheff{é} estimate to a robust two-point testing problem that leads to the construction of robust estimators adaptive to the proportion of contamination. Applying the general theory, we construct robust estimators for nonparametric density estimation, sparse linear regression and low-rank trace regression. We show that these new estimators achieve the minimax rate with optimal dependence on the contamination proportion. This testing procedure, Scheff{é} estimate, also enjoys an optimal rate in the exponent of the testing error, which may be of independent interest.

math.ST

Posterior Contraction Rates of the Phylogenetic Indian Buffet Processes

By expressing prior distributions as general stochastic processes, nonparametric Bayesian methods provide a flexible way to incorporate prior knowledge and constrain the latent structure in statistical inference. The Indian buffet process (IBP) is such an example that can be used to define a prior distribution on infinite binary features, where the exchangeability among subjects is assumed. The phylogenetic Indian buffet process (pIBP), a derivative of IBP, enables the modeling of non-exchangeability among subjects through a stochastic process on a rooted tree, which is similar to that used in phylogenetics, to describe relationships among the subjects. In this paper, we study the theoretical properties of IBP and pIBP under a binary factor model. We establish the posterior contraction rates for both IBP and pIBP and substantiate the theoretical results through simulation studies. This is the first work addressing the frequentist property of the posterior behaviors of IBP and pIBP. We also demonstrated its practical usefulness by applying pIBP prior to a real data example arising in the field of cancer genomics where the exchangeability among subjects is violated.

stat.ML

Change Point Analysis of Histone Modifications Reveals Epigenetic Blocks Linking to Physical Domains

Histone modification is a vital epigenetic mechanism for transcriptional control in eukaryotes. High-throughput techniques have enabled whole-genome analysis of histone modifications in recent years. However, most studies assume one combination of histone modification invariantly translates to one transcriptional output regardless of local chromatin environment. In this study we hypothesize that, the genome is organized into local domains that manifest similar enrichment pattern of histone modification, which leads to orchestrated regulation of expression of genes with relevant bio- logical functions. We propose a multivariate Bayesian Change Point (BCP) model to segment the Drosophila melanogaster genome into consecutive blocks on the basis of combinatorial patterns of histone marks. By modeling the sparse distribution of histone marks across the chromosome with a zero-inflated Gaussian mixture, our partitions capture local BLOCKs that manifest relatively homogeneous enrichment pattern of histone modifications. We further characterized BLOCKs by their transcription levels, distribution of genes, degree of co-regulation and GO enrichment. Our results demonstrate that these BLOCKs, although inferred merely from histone modifications, reveal strong relevance with physical domains, which suggest their important roles in chromatin organization and coordinated gene regulation.

q-bio.GN

Sparse CCA via Precision Adjusted Iterative Thresholding

Sparse Canonical Correlation Analysis (CCA) has received considerable attention in high-dimensional data analysis to study the relationship between two sets of random variables. However, there has been remarkably little theoretical statistical foundation on sparse CCA in high-dimensional settings despite active methodological and applied research activities. In this paper, we introduce an elementary sufficient and necessary characterization such that the solution of CCA is indeed sparse, propose a computationally efficient procedure, called CAPIT, to estimate the canonical directions, and show that the procedure is rate-optimal under various assumptions on nuisance parameters. The procedure is applied to a breast cancer dataset from The Cancer Genome Atlas project. We identify methylation probes that are associated with genes, which have been previously characterized as prognosis signatures of the metastasis of breast cancer.

math.ST

Asymptotically Normal and Efficient Estimation of Covariate-Adjusted Gaussian Graphical Model

A tuning-free procedure is proposed to estimate the covariate-adjusted Gaussian graphical model. For each finite subgraph, this estimator is asymptotically normal and efficient. As a consequence, a confidence interval can be obtained for each edge. The procedure enjoys easy implementation and efficient computation through parallel estimation on subgraphs or edges. We further apply the asymptotic normality result to perform support recovery through edge-wise adaptive thresholding. This support recovery procedure is called ANTAC, standing for Asymptotically Normal estimation with Thresholding after Adjusting Covariates. ANTAC outperforms other methodologies in the literature in a range of simulation studies. We apply ANTAC to identify gene-gene interactions using an eQTL dataset. Our result achieves better interpretability and accuracy in comparison with CAMPE.

stat.ME