SearcharxivSearch

arXiv subjects

Zixuan Cang

Publications and source records attributed to Zixuan Cang.

At least 19 recordsLinked to original sources

RAFT-UP: Robust Alignment for Spatial Transcriptomics with Explicit Control of Spatial Distortion

Spatial transcriptomics (ST) profiles gene expression across a tissue section while preserving the spatial coordinates. Because current ST technologies typically profile two-dimensional tissue slices, integrating and aligning slices from different regions of the same three-dimensional tissue or from samples under different conditions enables analyses that reveal 3D organization and condition-associated spatial patterns. Two major challenges remain. First, interpretable and flexible control over spatial distortion is needed because rigid transformations can be overly restrictive, whereas highly deformable mappings may arbitrarily distort spatial proximity. Second, biologically plausible matching is also needed, especially when the slices overlap partially. Here, we introduce RAFT-UP, a tool for robust ST alignment that provides explicit control over spatial distance preservation through a fused supervised Gromov-Wasserstein (FsGW) optimal transport framework. FsGW combines expression and spatial information, incorporates spot-wise constraints to discourage biologically implausible matches, and enforces a pairwise distance-consistency constraint that prevents mapping two pairs of spots when their spatial distances differ beyond a specified tolerance. We demonstrate that RAFT-UP accurately aligns slices from different regions of the same tissue and slices from different samples. Benchmarking shows that RAFT-UP improves spatial distance preservation while achieving spot label matching accuracy comparable to state-of-the-art methods. Finally, we demonstrate RAFT-UP on two spatially constrained downstream applications, including spatiotemporal mapping of developing mouse midbrain and comparative cross-slice analysis of cell-cell communication. RAFT-UP is available as open-source software.

q-bio.QM

Synchronization of Unbalanced Dynamical Optimal Transport across Multiple Spaces

Many biological systems are observed through heterogeneous modalities, requiring transport models that couple dynamics across spaces while allowing mass variation. To address this challenge, we introduce Unbalanced Synchronized Optimal Transport (UnSyncOT), a novel dynamical framework that synchronizes transport-reaction flows between spaces via either geometric embeddings (Monge type) or Markov kernels (Kantorovich type). For both cases we prove that UnSyncOT can be reduced to a single-space problem: the Monge model becomes a Benamou-Brenier problem with a metric-modified kinetic energy, and the Kantorovich model yields a nonlocal action induced by the synchronization operator, both of which fit within a dissipation-distance formulation. We also analyze the pure transport (Wasserstein) and pure reaction (Fisher-Rao) limits and derive structural properties. For the Kantorovich case we propose an approximate UnSyncOT by introducing a Hellinger-Kantorovich based trapezoidal time discretization of the secondary action for efficient computation. Finally we present staggered-grid discretizations and primal-dual solvers, validate the convergence, stability, and efficiency, and demonstrate coherent dynamics reconstructions across spaces.

math.OC

PHD-MS: Multiscale Domain Identification for Spatial Transcriptomics via Persistent Homology

Spatial transcriptomics (ST) measures gene expression at a set of spatial locations in a tissue. Communities of nearby cells that express similar genes form \textit{spatial domains}. Specialized ST clustering algorithms have been developed to identify these spatial domains. These methods often identify spatial domains at a single morphological scale, and interactions across multiple scales are often overlooked. For example, large cellular communities often contain smaller substructures, and heterogeneous frontier regions often lie between homogeneous domains. Topological data analysis (TDA) is an emerging mathematical toolkit that studies the underlying features of data at various geometric scales. It is especially useful for analyzing complex biological datasets with multiscale characteristics. Using TDA, we develop Persistent Homology for Domains at Multiple Scales (PHD-MS) to locate tissue structures that persist across morphological scales. We apply PHD-MS to highlight multiscale spatial domains in several tissue types and ST technologies. We also compare PHD-MS domains against ground-truth domains in expert-annotated tissues, where PHD-MS outperforms traditional clustering approaches. PHD-MS is available as an open-source software package with an interactive graphical user interface for exploring the identified multiscale domains.

q-bio.QM

OTMol: Robust Molecular Structure Comparison via Optimal Transport

Root-mean-square deviation (RMSD) is widely used to assess structural similarity in systems ranging from flexible ligand conformers to complex molecular cluster configurations. Despite its wide utility, RMSD calculation is often challenged by inconsistent atom ordering, indistinguishable configurations in molecular clusters, and potential chirality inversion during alignment. These issues highlight the necessity of accurate atom-to-atom correspondence as a prerequisite for meaningful alignment. Traditional approaches often rely on heuristic cost matrices combined with the Hungarian algorithm, yet these methods underutilize the rich intra-molecular structural information and may fail to generalize across chemically diverse systems. In this work, we introduce OTMol, a method that formulates the molecular alignment task as a fused supervised Gromov-Wasserstein (fsGW) optimal transport problem. By leveraging the intrinsic geometric and topological relationships within each molecule, OTMol eliminates the need for manually defined cost functions and enables a principled, data-driven matching strategy. Importantly, OTMol preserves key chemical features such as molecular chirality and bond connectivity consistency. We evaluate OTMol across a wide range of molecular systems, including Adenosine triphosphate, Imatinib, lipids, small peptides, and water clusters, and demonstrate that it consistently achieves low RMSD values while preserving computational efficiency. Importantly, OTMol maintains molecular integrity by enforcing one-to-one mappings between entire molecules, thereby avoiding erroneous many-to-one alignments that often arise in comparing molecular clusters. Our results underscore the utility of optimal transport theory for molecular alignment and offer a generalizable framework applicable to structural comparison tasks in cheminformatics, molecular modeling, and related disciplines.

q-bio.BM

Synchronized Optimal Transport for Joint Modeling of Dynamics Across Multiple Spaces

Optimal transport has been an essential tool for reconstructing dynamics from complex data. With the increasingly available multifaceted data, a system can often be characterized across multiple spaces. Therefore, it is crucial to maintain coherence in the dynamics across these diverse spaces. To address this challenge, we introduce Synchronized Optimal Transport (SyncOT), a novel approach to jointly model dynamics that represent the same system through multiple spaces. With given correspondence between the spaces, SyncOT minimizes the aggregated cost of the dynamics induced across all considered spaces. The problem is discretized into a finite-dimensional convex problem using a staggered grid. Primal-dual algorithm-based approaches are then developed to solve the discretized problem. Various numerical experiments demonstrate the capabilities and properties of SyncOT and validate the effectiveness of the proposed algorithms.

math.OC

Supervised Gromov-Wasserstein Optimal Transport

We introduce the supervised Gromov-Wasserstein (sGW) optimal transport, an extension of Gromov-Wasserstein by incorporating potential infinity patterns in the cost tensor. sGW enables the enforcement of application-induced constraints such as the preservation of pairwise distances by implementing the constraints as an infinity pattern. A numerical solver is proposed for the sGW problem and the effectiveness is demonstrated in various numerical experiments. The high-order constraints in sGW are transferred to constraints on the coupling matrix by solving a minimal vertex cover problem. The transformed problem is solved by the Mirror-C descent iteration coupled with the supervised optimal transport solver. In the numerical experiments, we first validate the proposed framework by applying it to matching synthetic datasets and investigating the impact of the model parameters. Additionally, we successfully apply sGW to real single-cell RNA sequencing data. Through comparisons with other Gromov-Wasserstein variants on real data, we demonstrate that sGW offers the novel utility of controlling distance preservation, leading to the automatic estimation of overlapping portions of datasets, which brings improved stability and flexibility in data-driven applications.

math.OC

Poisson-Boltzmann based machine learning (PBML) model for electrostatic analysis

Electrostatics is of paramount importance to chemistry, physics, biology, and medicine. The Poisson-Boltzmann (PB) theory is a primary model for electrostatic analysis. However, it is highly challenging to compute accurate PB electrostatic solvation free energies for macromolecules due to the nonlinearity, dielectric jumps, charge singularity , and geometric complexity associated with the PB equation. The present work introduces a PB based machine learning (PBML) model for biomolecular electrostatic analysis. Trained with the second-order accurate MIBPB solver, the proposed PBML model is found to be more accurate and faster than several eminent PB solvers in electrostatic analysis. The proposed PBML model can provide highly accurate PB electrostatic solvation free energy of new biomolecules or new conformations generated by molecular dynamics with much reduced computational cost.

physics.chem-ph

Topological and geometric analysis of cell states in single-cell transcriptomic data

Single-cell RNA sequencing (scRNA-seq) enables dissecting cellular heterogeneity in tissues, resulting in numerous biological discoveries. Various computational methods have been devised to delineate cell types by clustering scRNA-seq data where the clusters are often annotated using prior knowledge of marker genes. In addition to identifying pure cell types, several methods have been developed to identify cells undergoing state transitions which often rely on prior clustering results. Present computational approaches predominantly investigate the local and first-order structures of scRNA-seq data using graph representations, while scRNA-seq data frequently displays complex high-dimensional structures. Here, we present a tool, scGeom for exploiting the multiscale and multidimensional structures in scRNA-seq data by inspecting the geometry via graph curvature and topology via persistent homology of both cell networks and gene networks. We demonstrate the utility of these structural features for reflecting biological properties and functions in several applications where we show that curvatures and topological signatures of cell and gene networks can help indicate transition cells and developmental potency of cells. We additionally illustrate that the structural characteristics can improve the classification of cell types.

q-bio.QM

AVIDA: Alternating method for Visualizing and Integrating Data

High-dimensional multimodal data arises in many scientific fields. The integration of multimodal data becomes challenging when there is no known correspondence between the samples and the features of different datasets. To tackle this challenge, we introduce AVIDA, a framework for simultaneously performing data alignment and dimension reduction. In the numerical experiments, Gromov-Wasserstein optimal transport and t-distributed stochastic neighbor embedding are used as the alignment and dimension reduction modules respectively. We show that AVIDA correctly aligns high-dimensional datasets without common features with four synthesized datasets and two real multimodal single-cell datasets. Compared to several existing methods, we demonstrate that AVIDA better preserves structures of individual datasets, especially distinct local structures in the joint low-dimensional visualization, while achieving comparable alignment performance. Such a property is important in multimodal single-cell data analysis as some biological processes are uniquely captured by one of the datasets. In general applications, other methods can be used for the alignment and dimension reduction modules.

q-bio.QM

Supervised Optimal Transport

Optimal Transport, a theory for optimal allocation of resources, is widely used in various fields such as astrophysics, machine learning, and imaging science. However, many applications impose elementwise constraints on the transport plan which traditional optimal transport cannot enforce. Here we introduce Supervised Optimal Transport (sOT) that formulates a constrained optimal transport problem where couplings between certain elements are prohibited according to specific applications. sOT is proved to be equivalent to an $l^1$ penalized optimization problem, from which efficient algorithms are designed to solve its entropy regularized formulation. We demonstrate the capability of sOT by comparing it to other variants and extensions of traditional OT in color transfer problem. We also study the barycenter problem in sOT formulation, where we discover and prove a unique reverse and portion selection (control) mechanism. Supervised optimal transport is broadly applicable to applications in which constrained transport plan is involved and the original unit should be preserved by avoiding normalization.

math.OC

A review of mathematical representations of biomolecules

Recently, machine learning (ML) has established itself in various worldwide benchmarking competitions in computational biology, including Critical Assessment of Structure Prediction (CASP) and Drug Design Data Resource (D3R) Grand Challenges. However, the intricate structural complexity and high ML dimensionality of biomolecular datasets obstruct the efficient application of ML algorithms in the field. In addition to data and algorithm, an efficient ML machinery for biomolecular predictions must include structural representation as an indispensable component. Mathematical representations that simplify the biomolecular structural complexity and reduce ML dimensionality have emerged as a prime winner in D3R Grand Challenges. This review is devoted to the recent advances in developing low-dimensional and scalable mathematical representations of biomolecules in our laboratory. We discuss three classes of mathematical approaches, including algebraic topology, differential geometry, and graph theory. We elucidate how the physical and biological challenges have guided the evolution and development of these mathematical apparatuses for massive and diverse biomolecular data. We focus the performance analysis on the protein-ligand binding predictions in this review although these methods have had tremendous success in many other applications, such as protein classification, virtual screening, and the predictions of solubility, solvation free energy, toxicity, partition coefficient, protein folding stability changes upon mutation, etc.

q-bio.BM

Persistent cohomology for data with multicomponent heterogeneous information

Persistent homology is a powerful tool for characterizing the topology of a data set at various geometric scales. When applied to the description of molecular structures, persistent homology can capture the multiscale geometric features and reveal certain interaction patterns in terms of topological invariants. However, in addition to the geometric information, there is a wide variety of non-geometric information of molecular structures, such as element types, atomic partial charges, atomic pairwise interactions, and electrostatic potential function, that is not described by persistent homology. Although element specific homology and electrostatic persistent homology can encode some non-geometric information into geometry based topological invariants, it is desirable to have a mathematical framework to systematically embed both geometric and non-geometric information, i.e., multicomponent heterogeneous information, into unified topological descriptions. To this end, we propose a mathematical framework based on persistent cohomology. In our framework, non-geometric information can be either distributed globally or resided locally on the datasets in the geometric sense and can be properly defined on topological spaces, i.e., simplicial complexes. Using the proposed persistent cohomology based framework, enriched barcodes are extracted from datasets to represent heterogeneous information. We consider a variety of datasets to validate the present formulation and illustrate the usefulness of the proposed persistent cohomology. It is found that the proposed framework using cohomology boosts the performance of persistent homology based methods in the protein-ligand binding affinity prediction on massive biomolecular datasets.

q-bio.QM

Mathematical deep learning for pose and binding affinity prediction and ranking in D3R Grand Challenges

Advanced mathematics, such as multiscale weighted colored graph and element specific persistent homology, and machine learning including deep neural networks were integrated to construct mathematical deep learning models for pose and binding affinity prediction and ranking in the last two D3R grand challenges in computer-aided drug design and discovery. D3R Grand Challenge 2 (GC2) focused on the pose prediction and binding affinity ranking and free energy prediction for Farnesoid X receptor ligands. Our models obtained the top place in absolute free energy prediction for free energy Set 1 in Stage 2. The latest competition, D3R Grand Challenge 3 (GC3), is considered as the most difficult challenge so far. It has 5 subchallenges involving Cathepsin S and five other kinase targets, namely VEGFR2, JAK2, p38-$α$, TIE2, and ABL1. There is a total of 26 official competitive tasks for GC3. Our predictions were ranked 1st in 10 out of 26 official competitive tasks.

q-bio.BM

Evolutionary homology on coupled dynamical systems

Time dependence is a universal phenomenon in nature, and a variety of mathematical models in terms of dynamical systems have been developed to understand the time-dependent behavior of real-world problems. Originally constructed to analyze the topological persistence over spatial scales, persistent homology has rarely been devised for time evolution. We propose the use of a new filtration function for persistent homology which takes as input the adjacent oscillator trajectories of a dynamical system. We also regulate the dynamical system by a weighted graph Laplacian matrix derived from the network of interest, which embeds the topological connectivity of the network into the dynamical system. The resulting topological signatures, which we call evolutionary homology (EH) barcodes, reveal the topology-function relationship of the network and thus give rise to the quantitative analysis of nodal properties. The proposed EH is applied to protein residue networks for protein thermal fluctuation analysis, rendering the most accurate B-factor prediction of a set of 364 proteins. This work extends the utility of dynamical systems to the quantitative modeling and analysis of realistic physical systems.

math.AT

Representability of algebraic topology for biomolecules in machine learning based scoring and virtual screening

This work introduces a number of algebraic topology approaches, such as multicomponent persistent homology, multi-level persistent homology and electrostatic persistence for the representation, characterization, and description of small molecules and biomolecular complexes. Multicomponent persistent homology retains critical chemical and biological information during the topological simplification of biomolecular geometric complexity. Multi-level persistent homology enables a tailored topological description of inter- and/or intra-molecular interactions of interest. Electrostatic persistence incorporates partial charge information into topological invariants. These topological methods are paired with Wasserstein distance to characterize similarities between molecules and are further integrated with a variety of machine learning algorithms, including k-nearest neighbors, ensemble of trees, and deep convolutional neural networks, to manifest their descriptive and predictive powers for chemical and biological problems. Extensive numerical experiments involving more than 4,000 protein-ligand complexes from the PDBBind database and near 100,000 ligands and decoys in the DUD database are performed to test respectively the scoring power and the virtual screening power of the proposed topological approaches. It is demonstrated that the present approaches outperform the modern machine learning based methods in protein-ligand binding affinity predictions and ligand-decoy discrimination.

q-bio.QM

Analysis and prediction of protein folding energy changes upon mutation by element specific persistent homology

Motivation: Site directed mutagenesis is widely used to understand the structure and function of biomolecules. Computational prediction of protein mutation impacts offers a fast, economical and potentially accurate alternative to laboratory mutagenesis. Most existing methods rely on geometric descriptions, this work introduces a topology based approach to provide an entirely new representation of protein mutation impacts that could not be obtained from conventional techniques. Results: Topology based mutation predictor (T-MP) is introduced to dramatically reduce the geometric complexity and number of degrees of freedom of proteins, while element specific persistent homology is proposed to retain essential biological information. The present approach is found to outperform other existing methods in globular protein mutation impact predictions. A Pearson correlation coefficient of 0.82 with an RMSE of 0.92 kcal/mol is obtained on a test set of 350 mutation samples. For the prediction of membrane protein stability changes upon mutation, the proposed topological approach has a 84% higher Pearson correlation coefficient than the current state-of-the-art empirical methods, achieving a Pearson correlation of 0.57 and an RMSE of 1.09 kcal/mol in a 5-fold cross validation on a set of 223 membrane protein mutation samples.

q-bio.QM

Topological fingerprints reveal protein-ligand binding mechanism

Protein-ligand binding is a fundamental biological process that is paramount to many other biological processes, such as signal transduction, metabolic pathways, enzyme construction, cell secretion, gene expression, etc. Accurate prediction of protein-ligand binding affinities is vital to rational drug design and the understanding of protein-ligand binding and binding induced function. Existing binding affinity prediction methods are inundated with geometric detail and involve excessively high dimensions, which undermines their predictive power for massive binding data. Topology provides an ultimate level of abstraction and thus incurs too much reduction in geometric information. Persistent homology embeds geometric information into topological invariants and bridges the gap between complex geometry and abstract topology. However, it over simplifies biological information. This work introduces element specific persistent homology (ESPH) to retain crucial biological information during topological simplification. The combination of ESPH and machine learning gives rise to one of the most efficient and powerful tools for revealing protein-ligand binding mechanism and for predicting binding affinities.

q-bio.QM

TopologyNet: Topology based deep convolutional neural networks for biomolecular property predictions

Although deep learning approaches have had tremendous success in image, video and audio processing, computer vision, and speech recognition, their applications to three-dimensional (3D) biomolecular structural data sets have been hindered by the entangled geometric complexity and biological complexity. We introduce topology, i.e., element specific persistent homology (ESPH), to untangle geometric complexity and biological complexity. ESPH represents 3D complex geometry by one-dimensional (1D) topological invariants and retains crucial biological information via a multichannel image representation. It is able to reveal hidden structure-function relationships in biomolecules. We further integrate ESPH and convolutional neural networks to construct a multichannel topological neural network (TopologyNet) for the predictions of protein-ligand binding affinities and protein stability changes upon mutation. To overcome the limitations to deep learning arising from small and noisy training sets, we present a multitask topological convolutional neural network (MT-TCNN). We demonstrate that the present TopologyNet architectures outperform other state-of-the-art methods in the predictions of protein-ligand binding affinities, globular protein mutation impacts, and membrane protein mutation impacts.

q-bio.QM