SearcharxivSearch

arXiv subjects

Xiaoqi Wei

Publications and source records attributed to Xiaoqi Wei.

9 recordsLinked to original sources

Synchronization of Unbalanced Dynamical Optimal Transport across Multiple Spaces

Many biological systems are observed through heterogeneous modalities, requiring transport models that couple dynamics across spaces while allowing mass variation. To address this challenge, we introduce Unbalanced Synchronized Optimal Transport (UnSyncOT), a novel dynamical framework that synchronizes transport-reaction flows between spaces via either geometric embeddings (Monge type) or Markov kernels (Kantorovich type). For both cases we prove that UnSyncOT can be reduced to a single-space problem: the Monge model becomes a Benamou-Brenier problem with a metric-modified kinetic energy, and the Kantorovich model yields a nonlocal action induced by the synchronization operator, both of which fit within a dissipation-distance formulation. We also analyze the pure transport (Wasserstein) and pure reaction (Fisher-Rao) limits and derive structural properties. For the Kantorovich case we propose an approximate UnSyncOT by introducing a Hellinger-Kantorovich based trapezoidal time discretization of the secondary action for efficient computation. Finally we present staggered-grid discretizations and primal-dual solvers, validate the convergence, stability, and efficiency, and demonstrate coherent dynamics reconstructions across spaces.

math.OC

Computational Drug Repurposing for Alzheimer's Disease via Sheaf Theoretic Population-Scale Analysis of snRNA-seq Data

Single-cell and single-nucleus RNA sequencing (scRNA-seq /snRNA-seq) are widely used to reveal heterogeneity in cells, showing a growing potential for precision and personalized medicine. Nonetheless, sustainable drug discovery must be based on a population-level understanding of molecular mechanisms, which calls for the population-scale analysis of scRNA-seq/snRNA-seq data. This work introduces a sequential target-drug selection model for drug repurposing against Alzheimer's Disease (AD) targets inferred from population-level snRNA-seq studies of AD progression in microglia cells as well as different cell types taken from an AD affected brain vascular tissue atlas, involving hundreds of thousands of nuclei from multi-patient and multi-regional studies. We utilize Persistent Sheaf Laplacians (PSL) to facilitate a Protein-Protein Interaction (PPI) analysis of AD targets inferred from differential gene expression (DEG), and then use machine learning models to predict repurpose-able DrugBank compounds for molecular targeting. We screen the efficacy of different DrugBank small compounds and further examine their central nervous system (CNS)-relevant ADMET (Absorption, Distribution, Metabolism, Excretion, and Toxicity), resulting in a list of lead candidates for AD treatment. The list of significant genes establishes a target domain for effective machine learning based AD drug repurposing analysis of DrugBank small compounds to treat AD related molecular targets.

q-bio.MN

OTMol: Robust Molecular Structure Comparison via Optimal Transport

Root-mean-square deviation (RMSD) is widely used to assess structural similarity in systems ranging from flexible ligand conformers to complex molecular cluster configurations. Despite its wide utility, RMSD calculation is often challenged by inconsistent atom ordering, indistinguishable configurations in molecular clusters, and potential chirality inversion during alignment. These issues highlight the necessity of accurate atom-to-atom correspondence as a prerequisite for meaningful alignment. Traditional approaches often rely on heuristic cost matrices combined with the Hungarian algorithm, yet these methods underutilize the rich intra-molecular structural information and may fail to generalize across chemically diverse systems. In this work, we introduce OTMol, a method that formulates the molecular alignment task as a fused supervised Gromov-Wasserstein (fsGW) optimal transport problem. By leveraging the intrinsic geometric and topological relationships within each molecule, OTMol eliminates the need for manually defined cost functions and enables a principled, data-driven matching strategy. Importantly, OTMol preserves key chemical features such as molecular chirality and bond connectivity consistency. We evaluate OTMol across a wide range of molecular systems, including Adenosine triphosphate, Imatinib, lipids, small peptides, and water clusters, and demonstrate that it consistently achieves low RMSD values while preserving computational efficiency. Importantly, OTMol maintains molecular integrity by enforcing one-to-one mappings between entire molecules, thereby avoiding erroneous many-to-one alignments that often arise in comparing molecular clusters. Our results underscore the utility of optimal transport theory for molecular alignment and offer a generalizable framework applicable to structural comparison tasks in cheminformatics, molecular modeling, and related disciplines.

q-bio.BM

Persistent Sheaf Laplacian Analysis of Protein Flexibility

Protein flexibility, measured by the B-factor or Debye-Waller factor, is essential for protein functions such as structural support, enzyme activity, cellular communication, and molecular transport. Theoretical analysis and prediction of protein flexibility are crucial for protein design, engineering, and drug discovery. In this work, we introduce the persistent sheaf Laplacian (PSL), an effective tool in topological data analysis, to model and analyze protein flexibility. By representing the local topology and geometry of protein atoms through the multiscale harmonic and non-harmonic spectra of PSLs, the proposed model effectively captures protein flexibility and provides accurate, robust predictions of protein B-factors. Our PSL model demonstrates an increase in accuracy of 32% compared to the classical Gaussian network model (GNM) in predicting B-factors for a dataset of 364 proteins. Additionally, we construct a blind machine learning prediction method utilizing global and local protein features. Extensive computations and comparisons validate the effectiveness of the proposed PSL model for B-factor predictions.

q-bio.BM

Persistent Topological Laplacians -- a Survey

Persistent topological Laplacians constitute a new class of tools in topological data analysis (TDA). They are motivated by the necessity to address challenges encountered in persistent homology when handling complex data. These Laplacians combines multiscale analysis with topological techniques to characterize the topological and geometrical features of functions and data. Their kernels fully retrieve the topological invariants of corresponding persistent homology, while their non-harmonic spectra provide supplementary information. Persistent topological Laplacians have demonstrated superior performance over persistent homology in analyzing large-scale protein engineering datasets. In this survey, we offer a pedagogical review of persistent topological Laplacians formulated in various mathematical settings, including simplicial complexes, path complexes, flag complexes, digraphs, hypergraphs, hyperdigraphs, cellular sheaves, as well as $N$-chain complexes.

math.AT

Persistent sheaf Laplacians

Recently various types of topological Laplacians have been studied from the perspective of data analysis. The spectral theory of these Laplacians has significantly extended the scope of algebraic topology and data analysis. Inspired by the theory of persistent Laplacians and cellular sheaves, this work develops the theory of persistent sheaf Laplacians for cellular sheaves, and describes how to construct sheaves for a point cloud where each point is associated with a quantity that can be devised to embed physical properties. As a result, the spectra of persistent sheaf Laplacians encode both geometrical and non-geometrical information of the given point cloud. The theory of persistent sheaf Laplacians is an elegant method for fusing different types of data and has huge potential for future development.

math.AT

Persistent topological Laplacian analysis of SARS-CoV-2 variants

Topological data analysis (TDA) is an emerging field in mathematics and data science. Its central technique, persistent homology, has had tremendous success in many science and engineering disciplines. However, persistent homology has limitations, including its incapability of describing the homotopic shape evolution of data during filtration. Persistent topological Laplacians (PTLs), such as persistent Laplacian and persistent sheaf Laplacian, were proposed to overcome the drawback of persistent homology. In this work, we examine the modeling and analysis power of PTLs in the study of the protein structures of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) spike receptor binding domain (RBD) and its variants, i.e., Alpha, Beta, Gamma, BA.1, and BA.2. First, we employ PTLs to study how the RBD mutation-induced structural changes of RBD-angiotensin-converting enzyme 2 (ACE2) binding complexes are captured in the changes of spectra of the PTLs among SARS-CoV-2 variants. Additionally, we use PTLs to analyze the binding of RBD and ACE2-induced structural changes of various SARS-CoV-2 variants. Finally, we explore the impacts of computationally generated RBD structures on PTL-based machine learning, including deep learning, and predictions of deep mutational scanning datasets for the SARS-CoV-2 Omicron BA.2 variant. Our results indicate that PTLs have advantages over persistent homology in analyzing protein structural changes and provide a powerful new TDA tool for data science.

q-bio.QM

Emerging dominant SARS-CoV-2 variants

Accurate and reliable forecasting of emerging dominant severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) variants enables policymakers and vaccine makers to get prepared for future waves of infections. The last three waves of SARS-CoV-2 infections caused by dominant variants Omicron (BA.1), BA.2, and BA.4/BA.5 were accurately foretold by our artificial intelligence (AI) models built with biophysics, genotyping of viral genomes, experimental data, algebraic topology, and deep learning. Based on newly available experimental data, we analyzed the impacts of all possible viral spike (S) protein receptor-binding domain (RBD) mutations on the SARS-CoV-2 infectivity. Our analysis sheds light on viral evolutionary mechanisms, i.e., natural selection through infectivity strengthening and antibody resistance. We forecast that BA.2.10.4, BA.2.75, BQ.1.1, and particularly, BA.2.75+R346T, have high potential to become new dominant variants to drive the next surge.

q-bio.PE

The edge ideals of the join of some vertex weighted oriented graphs

In this paper, we describe primary decomposition of the edge ideal of the join of some graphs in terms of that information of the edge ideal of every weighted oriented graph. Meanwhile, we also study depth and regularity of symbolic powers and ordinary powers of such an edge ideal. We explicitly compute depth and regularity of ordinary powers of the edge ideal of the join of two graphs consisting of isolated vertices, and also provide upper bounds of regularity of symbolic powers of such an edge ideal. For the edge ideal of the join of two graphs with at least an oriented edge for per graph, we give the exact formulas for their depth and regularity, and also provide the upper bounds of regularity of ordinary powers of such an edge ideal. Some examples show that these upper bounds can be obtained, but may be strict.

math.AC