SearcharxivSearch

arXiv subjects

Guo-Wei Wei

Publications and source records attributed to Guo-Wei Wei.

At least 19 recordsLinked to original sources

GrassTop: Grassmannian k-mer Topology for Viral Classification and Phylogenetic Analysis

We introduce GrassTop, a genome representation that integrates Grassmann manifolds and algebraic topology for viral classification and phylogenetic analysis. The framework begins by constructing multiscale topological and spectral descriptors of (k)-mer positional patterns. It then extracts a low-rank subspace that summarizes variation across the filtration and compares genomes using a Grassmannian distance. Although the reported implementation uses the chordal distance, the framework is not restricted to this particular choice. We evaluate GrassTop on four families of viral classification datasets, four phylogenetic clustering datasets, and a sequence perturbation experiment. Under the reported 5-nearest-neighbor protocol, GrassTop achieves higher scores than five published alignment-free reference methods across all reported classification metrics and datasets. Its UPGMA (unweighted pair-group method using arithmetic averages) trees achieve an average label purity of 1.0 on every phylogenetic dataset. The perturbation experiment provides a more nuanced result: the subspace representation differs most clearly from direct comparison of the unprojected feature matrices for SARS-CoV-2, whereas the differences are smaller or non-monotonic for the other datasets. Overall, these results support GrassTop as an effective topological-geometric representation for viral classification and phylogenetic analysis.

q-bio.PE

Data-Adaptive Grassmann Manifold Representations for Spatial Transcriptomics Alignment

Spatial transcriptomics measures gene expression together with spatial coordinates, but many existing analysis methods represent each spot primarily by a single feature vector. We propose GrassST, a subspace method for spatial transcriptomics analysis and cross-slice alignment. For each spatial spot, GrassST constructs a neighborhood from tissue coordinates, fits a low-dimensional subspace to the embedded expression profiles in that neighborhood, and represents the spot by the resulting subspace. These representations are points on a Grassmann manifold and can be compared using distances between subspaces. GrassST selects the neighborhood size and subspace rank from the spectral energy of the data, allowing these parameters to vary across datasets rather than being fixed globally. Experiments on four spatial transcriptomics datasets show that GrassST achieves competitive clustering and cross-slice integration performance under a unified evaluation pipeline.

math.AG

Localized Persistent Commutative Algebra

We develop a localized persistent theory of commutative algebra for Stanley-Reisner rings, based on local cohomology supported at a coordinate prime rather than at the maximal ideal. The construction is modeled on the persistent Stanley-Reisner theory of Suwayyid and Wei (arXiv:2503.23482) and its functorial development for graphs and hypergraphs (arXiv:2512.17619), in which invariants of the face ring such as graded Betti numbers and f- and h-vectors are persisted across a filtration. That framework is built from the minimal free resolution and is thus Tor-theoretic; we work instead on the injective side, and the resulting modules record information localized at a single vertex, complementing the global picture given by maximal-support local cohomology. For a vertex prime $p_i = (x_j : j \neq i)$ we prove an exact $\mathbb{Z}^n$-graded decomposition of $H^q_{p_i}(k[Δ])$ into the maximal-support local cohomology of the deletion and of the link of the vertex $i$, the first in $x_i$-degree zero and the second repeated in every positive $x_i$-degree; at the level of graded dimensions this recovers the vertex-prime case of Rahimi's bigraded formula. With Hochster's formula this yields a closed combinatorial description of every multigraded piece. Building on this structure we introduce per-vertex persistent local cohomology numbers, prove a persistent links-Hochster formula, obtain interval decompositions of the resulting reversed-arrow persistence modules and a bottleneck stability theorem, retain multiplication by the uninverted variable as a morphism of persistence modules that the two barcodes alone do not determine, and extend the theory to an arbitrary coordinate prime, where the multiplication maps of the uninverted variables assemble into a commuting Boolean diagram of persistence modules.

math.AC

Parametrization of Symmetry in Data

Symmetry plays a fundamental role in understanding natural phenomena and mathematical structures. This work develops a comprehensive theory for studying the persistent symmetries and degree of asymmetry of finite point configurations over parameterization in metric spaces. Leveraging category theory and span categories, we define persistent symmetry groups and introduce novel invariants called symmetry barcodes and polybarcodes that capture the birth, death, persistence, and reappearance of symmetries over parameter evolution. Metrics and stability theorems are established for these invariants. The concept of symmetry types is formalized via the action of isometry groups in configuration spaces. To quantitatively characterize symmetry and asymmetricity, measures such as degree of symmetry and symmetry defect are introduced, the latter revealing connections to approximate group theory in Euclidean settings. Moreover, a theory of persistence representations of persistence groups is developed, generalizing the classical decomposition theorem of persistence modules. Persistent Fourier analysis on persistence groups is further proposed to characterize dynamic phenomena including symmetry breaking and phase transitions. Algorithms for computing symmetry groups, barcodes, and symmetry defect in low-dimensional spaces are presented, complemented by discussions on extending symmetry analysis beyond geometric contexts. This work thus bridges geometric group theory, topological data analysis, representation theory, and machine learning, providing novel tools for the analysis of the parametrized symmetry of data.

math.AT

PSLL: Persistent Sheaf Laplacian Learning for Protein-Ligand Binding Affinity Prediction

Accurate prediction of protein-ligand binding affinity remains a central challenge in computational drug discovery due to the complex interplay among molecular geometry, physicochemical interactions, and atom-specific charge information. In this work, we introduce a Persistent Sheaf Laplacian learning (PSLL) framework for protein-ligand binding affinity prediction. The proposed approach constructs multiscale topological representations from three-dimensional protein-ligand complexes by incorporating atomic partial charges into sheaf restriction maps over Vietoris-Rips and alpha complex filtrations. To capture chemically diverse protein-ligand interactions, we introduce element-specific and category-specific atom-pair representations within the PSLL framework. Harmonic and non-harmonic spectra extracted from the resulting persistent sheaf Laplacians are used as molecular descriptors. To complement the PSLL-derived molecular representation, we incorporate transformer-based protein embeddings and SMILES-derived ligand descriptors for binding affinity prediction. The scoring power of the proposed multiscale PSLL model is validated against existing state-of-the-art methods on three widely used PDBbind benchmark datasets, including PDBbind-v2007, PDBbind-v2013, and PDBbind-v2016. The computational results indicate that the proposed PSLL model achieves strong predictive performance across benchmark datasets, highlighting its potential as an interpretable and mathematically grounded framework with promising generalizability for molecular machine learning and drug discovery.

q-bio.BM

CAKR: Commutative algebra k-mer representations for genomics

Despite the availability of various sequence analysis models, comparative genomic analysis remains a challenge in genomics, genetics, and phylogenetics. Commutative algebra, a fundamental tool in algebraic geometry and number theory, has rarely been used in data and biological sciences. In this study, we introduce commutative algebra $k$-mer representations as a nonlinear algebraic framework for analyzing genomic sequences. This representation bridges commutative algebra, algebraic topology, combinatorics, and machine learning to establish a mathematical framework for comparative genomic analysis. We evaluate its effectiveness on three tasks including genetic variant classification, phylogenetic tree reconstruction, and viral classification, typically requiring alignment-based, alignment-free, and machine-learning approaches, respectively. In this work, we show that commutative algebra k-mer representations outperform five state-of-the-art sequence analysis methods across twelve primary datasets, with two additional supplementary fragment-placement benchmarks, especially in viral classification, and maintain relatively stable predictive accuracy as dataset size increases, underscoring scalability and robustness.

q-bio.QM

Weighted Hodge Laplacians on Manifolds with Boundary

The spectrum of the Hodge Laplacian on differential manifolds encodes rich topological and geometric information and thus provides a powerful tool for analyzing data on manifolds. However, the classical unweighted formulation is restricted in its ability to study data with varying local features. To address this limitation, we propose a weighted Hodge Laplacian framework for manifolds with boundary, both in theory and in computation, by incorporating a weight function on the manifold. Under appropriate boundary conditions, we formulate the corresponding weighted de Rham-Hodge theory, in which the kernel of the weighted Hodge Laplacian coincides with the weighted harmonic space, and remains isomorphic to the de Rham cohomology of the underlying manifold. The harmonic spectrum of the weighted Hodge Laplacian captures the global topological information, while its non-harmonic spectrum encodes the local geometric property induced by the weight. The proposed framework therefore enables the study of topological and geometric features of data on manifolds across varying weights, and in addition, allows local structure to be highlighted by choosing weights that emphasize regions of interest. We demonstrate the effectiveness of the proposed method through proof-of-principle experiments in protein flexibility analysis, and the results show its promise.

math.DG

Persistent Manifold Learning of Protein Properties

Predicting how tightly two biomolecules bind remains a major challenge, in part because different interaction classes present dissimilar interfaces, from compact metal-coordinated pockets to broad, featureless protein surfaces. We introduce persistent manifold learning (PML), a novel computational framework that describes a binding interface as a family of multiscale manifolds. Boundary-Induced Graph Laplacian, a discrete realization of de Rham-Hodge theory, then extracts topological invariants together with nonharmonic spectral information, capturing the geometry of an interface as well as its topology. These manifold embeddings are combined with protein and molecular language model representations and paired with gradient boosting decision trees. Our PML outperforms state-of-the-art methods on metalloprotein-ligand and protein-protein benchmarks.

q-bio.BM

AlphaFunctor: Bridging The Gap Between Protein Function Annotation and Property Prediction

The fundamental relationship among protein sequence, structure, function, and physicochemical properties is a central principle in biology. While in principle protein function and properties should be able to be derived directly from protein sequence, in practice protein function and property prediction methods have been designed around specific datasets and specific property or function subsets, leading to an enormous gap between function annotation and property prediction. To address these challenges, we introduce AlphaFunctor, a category theory based foundation model-like platform to bridge the gap between protein function annotation and property prediction. Based on the hypothesis that protein function and properties can be directly derived from protein sequence, AlphaFunctor predicts protein functions as represented by Gene Ontology terms directly from sequence. Using these function predictions, AlphaFunctor further maps protein functions using topological spectral theory, path-complex neural networks, and protein domain analysis onto downstream property prediction. AlphaFunctor is (pre)trained in nearly 0.6 million protein function data points to deliver the state-of-the-art protein function annotation on three benchmark datasets. Without task-specific network redesign, AlphaFunctor maps qualitative protein function annotation to various qualitative and quantitative protein property predictions, outperforming other dataset-specific and task-specific competing predictors.

q-bio.BM

Mayer Path Homology

We introduce Mayer path homology, a new homology theory for directed path complexes obtained by equipping path complexes with an $N$-nilpotent differential. The main novelty of this work is the introduction of an $N$-differential on path complexes, giving rise to $N$-chain complexes of $\partial$-invariant paths and Mayer path homology groups $H_n^{N,q}(P)$. We prove that this construction defines a canonical invariant of directed graphs and is more sensitive than standard path homology, distinguishing directed network motifs that ordinary path homology cannot separate. We further establish a complete classification of generators of $Ω_2^N$ and $Ω_3^N$, determining all admissible combinatorial types. Finally, we characterize elements of the first Mayer path cycles group $Z_1^{N,q}$ in terms of weighted directed cycles arising from spanning-tree constructions. These results provide the first systematic structural theory for Mayer path complexes and reveal new higher-order algebraic structures in directed graphs.

math.AT

VARIANT: Web Server for Decoding and Analyzing Viral Mutations at Genome and Protein Levels

A comprehensive analysis of viral mutations is essential for understanding viral evolution, disease epidemiology, diagnosis, drug resistance, etc. However, challenges remain in capturing complex mutation patterns and supporting diverse viral families with varying genome architectures. To address these needs, we present VARIANT, an web server for mutational analysis of RNA viral genomes and associated viral products across both single- and multi-segment virus genomes. The server takes as input a viral reference genome, a reference protein sequence, and/or multiple sequence alignment, and automatically provides full annotation of mutation types, including standard categories such as point mutations (missense, silent, and nonsense), insertions, deletions, or frameshift events in both coding and non-coding regions. In addition, VARIANT detects three biologically significant mutation patterns that are overlooked by conventional software/packages: ``row mutations'' (consecutive substitutions within a window of 3 nts), ``hot mutations'' (two non-consecutive substitutions within a window of 3 nts), and potential programmed ribosomal frameshifting (PRF) regions. The server currently contains automatic analysis of major viral pathogens, including SARS-CoV-2, HIV-1, Influenza H3N2, Ebola virus, and Chikungunya virus. It also allows users to analyze customized viruses. Users can track VARIANT analysis progress in real time, visualize mutation distributions, and download structured results in ZIP format. VARIANT also incorporates dual graph topology analysis to classify frameshifting element structures from dot-bracket notation input. This feature enables systematic comparison of RNA secondary structure motifs across viral families by mapping structures to a comprehensive library of dual graph topologies. The web server is freely available at https://variant.up.railway.app.

q-bio.QM

Subspace Tensor Orthogonal Rotation Model (STORM) for Batch Alignment, Cell Type Deconvolution, and Gene Imputation in Spatial Transcriptomic Data

Spatial transcriptomics data analysis integrates cellular transcriptional activity with spatial coordinates to identify spatial domains, infer cell-type dynamics, and characterize gene expression patterns within tissues. Despite recent advances, significant challenges remain, including the treatment of batch effects, the handling of mixed cell-type signals, and the imputation of poorly measured or missing gene expression. This work addresses these challenges by introducing a novel Subspace Tensor Orthogonal Rotation Model (STORM) that aligns multiple slices which vary in their spatial dimensions and geometry by considering them at the level of physical patterns or microenvironments. To this end, STORM presents an irregular tensor factorization technique for decomposing a collection of gene expression matrices and integrating them into a shared latent space for downstream analysis. In contrast to black-box deep learning approaches, the proposed model is inherently interpretable. Numerical experiments demonstrate state-of-the-art performance in vertical and horizontal batch integration, cell-type deconvolution, and unmeasured gene imputation for spatial transcriptomics data.

q-bio.QM

PETLS: PErsistent Topological Laplacian Software

Persistent topological Laplacians are operators that provide persistent Betti numbers and additional multiscale geometric information through the eigenvalues of the persistent topological Laplacian matrix. We introduce a framework and novel algorithm to aid in the computation of persistent topological Laplacians. We implement existing and new persistent Laplacian algorithms in an efficient and flexible C++ library with Python bindings, titled PETLS: PErsistent Topological Laplacian Software. As part of this library, we interface with several complexes commonly used in topological data analysis (TDA), such as simplicial, alpha, directed flag, Dowker, and cellular Sheaf. Because increased efficiency broadens the set of computationally feasible applications, we provide recommendations on how to use algorithms and complexes for data analysis in machine learning.

math.AT

Multi-dimensional Persistent Sheaf Laplacians for Image Analysis

We propose a multi-dimensional persistent sheaf Laplacian (MPSL) framework on simplicial complexes for image analysis. The proposed method is motivated by the strong sensitivity of commonly used dimensionality reduction techniques, such as principal component analysis (PCA), to the choice of reduced dimension. Rather than selecting a single reduced dimension or averaging results across dimensions, we exploit complementary advantages of multiple reduced dimensions. At a given dimension, image samples are regarded as simplicial complexes, and persistent sheaf Laplacians are utilized to extract a multiscale localized topological spectral representation for individual image samples. Statistical summaries of the resulting spectra are then aggregated across scales and dimensions to form multiscale multi-dimensional image representations. We evaluate the proposed framework on the COIL20 and ETH80 image datasets using standard classification protocols. Experimental results show that the proposed method provides more stable performance across a wide range of reduced dimensions and achieves consistent improvements to PCA-based baselines in moderate dimensional regimes.

cs.CV

Persistent Sheaf Laplacian Analysis of Protein Stability and Solubility Changes upon Mutation

Genetic mutations frequently disrupt protein structure, stability, and solubility, acting as primary drivers for a wide spectrum of diseases. Despite the critical importance of these molecular alterations, existing computational models often lack interpretability, and fail to integrate essential physicochemical interaction. To overcome these limitations, we propose SheafLapNet, a unified predictive framework grounded in the mathematical theory of Topological Deep Learning (TDL) and Persistent Sheaf Laplacian (PSL). Unlike standard Topological Data Analysis (TDA) tools such as persistent homology, which are often insensitive to heterogeneous information, PSL explicitly encodes specific physical and chemical information such as partial charges directly into the topological analysis. SheafLapNet synergizes these sheaf-theoretic invariants with advanced protein transformer features and auxiliary physical descriptors to capture intrinsic molecular interactions in a multiscale and mechanistic manner. To validate our framework, we employ rigorous benchmarks for both regression and classification tasks. For stability prediction, we utilize the comprehensive S2648 and S350 datasets. For solubility prediction, we employ the PON-Sol2 dataset, which provides annotations for increased, decreased, or neutral solubility changes. By integrating these multi-perspective features, SheafLapNet achieves state-of-the-art performance across these diverse benchmarks, demonstrating that sheaf-theoretic modeling significantly enhances both interpretability and generalizability in predicting mutation-induced structural and functional changes.

math.SP

Persistent commutative algebra on graphs and hypergraphs

We introduce a persistent commutative algebra for studying the algebraic and combinatorial evolution of edge ideals of graphs and hypergraphs under filtration. Building on the Persistent Stanley--Reisner Theory (PSRT), we develop the notion of persistent edge ideals and analyze their graded Betti numbers across the filtration of graphs or hypergraphs. To enable this analysis, we establish a persistent extension of Hochster's formula, providing a functorial correspondence between algebraic and topological persistence. We further examine the behavior of Betti splittings in the persistent setting, proving a general inequality that extends the classical splitting result to the filtration of monomial ideals. Motivated by graph-theoretic interpretations, we introduce persistent minimal vertex covers, which encode the temporal structure of combinatorial dependencies within evolving graphs or hypergraphs. Applications to alignment-free genomic classification and molecular isomer discrimination demonstrate the interpretability and representatbility of persistent edge ideals as algebraic invariants, bridging combinatorial commutative algebra and data science.

math.AC

Multiscale Grassmann Manifolds for Single-Cell Data Analysis

Single-cell data analysis seeks to characterize cellular heterogeneity based on high-dimensional gene expression profiles. Conventional approaches represent each cell as a vector in Euclidean space, which limits their ability to capture intrinsic correlations and multiscale geometric structures. We propose a multiscale framework based on Grassmann manifolds that integrates machine learning with subspace geometry for single-cell data analysis. By generating embeddings under multiple representation scales, the framework combines their features from different geometric views into a unified Grassmann manifold. A power-based scale sampling function is introduced to control the selection of scales and balance in- formation across resolutions. Experiments on nine benchmark single-cell RNA-seq datasets demonstrate that the proposed approach effectively preserves meaningful structures and provides stable clustering performance, particularly for small to medium-sized datasets. These results suggest that Grassmann manifolds offer a coherent and informative foundation for analyzing single cell data.

cs.LG

Commutative Algebra Modeling in Materials Science -- A Case Study on Metal-Organic Frameworks (MOFs)

Metal-organic frameworks (MOFs) are a class of important crystalline and highly porous materials whose hierarchical geometry and chemistry hinder interpretable predictions in materials properties. Commutative algebra is a branch of abstract algebra that has been rarely applied in data and material sciences. We introduce the first ever commutative algebra modeling and prediction in materials science. Specifically, category-specific commutative algebra (CSCA) is proposed as a new framework for MOF representation and learning. It integrates element-based categorization with multiscale algebraic invariants to encode both local coordination motifs and global network organization of MOFs. These algebraically consistent, chemically aware representations enable compact, interpretable, and data efficient modeling of MOF properties such as Henry's constants and uptake capacities for common gases. Compared to traditional geometric and graph-based approaches, CSCA achieves comparable or superior predictive accuracy while substantially improving interpretability and stability across data sets. By aligning commutative algebra with the chemical hierarchy, the CSCA establishes a rigorous and generalizable paradigm for understanding structure and property relationships in porous materials and provides a nonlinear algebra-based framework for data-driven material discovery.

cond-mat.mtrl-sci