SearcharxivSearch

arXiv subjects

Guowei Wei

Publications and source records attributed to Guowei Wei.

12 recordsLinked to original sources

Interpretability and Representability of Commutative Algebra, Algebraic Topology, and Topological Spectral Theory for Real-World Data

Recent years have witnessed a fast growth in mathematical artificial intelligence (AI). One of the most successful mathematical AI approaches is topological data analysis (TDA) via persistent homology (PH) that provides explainable AI (xAI) by extracting multiscale structural features from complex datasets. This work investigates the interpretability and representability of three foundational mathematical AI methods, PH, persistent Laplacians (PL) derived from spectral theory, and persistent commutative algebra (PCA) rooted in Stanley-Reisner theory. We apply these methods to a set of data, including geometric shapes, synthetic complexes, fullerene structures, and biomolecular systems to examine their geometric, topological and algebraic properties. PH captures topological invariants such as connected components, loops, and voids through persistence barcodes. PL extends PH by incorporating spectral information, quantifying topological invariants, geometric stiffness and connectivity via harmonic and non-harmonic spectra. PCA introduces algebraic invariants such as graded Betti numbers, facet persistence, and f/h-vectors, offering combinatorial, topological, geometric, and algebraic perspectives on data over scales. Comparative analysis reveals that while PH offers computational efficiency and intuitive visualization, PL provides enhanced geometric sensitivity, and PCA delivers rich algebraic interpretability. Together, these methods form a hierarchy of mathematical representations, enabling explainable and generalizable AI for real-world data.

math.GN

Persistent Directed Flag Laplacian

Topological data analysis (TDA) has had enormous success in science and engineering in the past decade. Persistent topological Laplacians (PTLs) overcome some limitations of persistent homology, a key technique in TDA, and provide substantial insight to the behavior of various geometric and topological objects. This work extends PTLs to directed flag complexes, which are an exciting generalization to flag complexes, also known as clique complexes, that arise naturally in many situations. We introduce the directed flag Laplacian and show that the proposed persistent directed flag Laplacian (PDFL) is a distinct way of analyzing these flag complexes. Example calculations are provided to demonstrate the potential of the proposed PDFL in real world applications.

math.AT

PLPCA: Persistent Laplacian Enhanced-PCA for Microarray Data Analysis

Over the years, Principal Component Analysis (PCA) has served as the baseline approach for dimensionality reduction in gene expression data analysis. It primary objective is to identify a subset of disease-causing genes from a vast pool of thousands of genes. However, PCA possesses inherent limitations that hinder its interpretability, introduce classification ambiguity, and fail to capture complex geometric structures in the data. Although these limitations have been partially addressed in the literature by incorporating various regularizers such as graph Laplacian regularization, existing improved PCA methods still face challenges related to multiscale analysis and capturing higher-order interactions in the data. To address these challenges, we propose a novel approach called Persistent Laplacian-enhanced Principal Component Analysis (PLPCA). PLPCA amalgamates the advantages of earlier regularized PCA methods with persistent spectral graph theory, specifically persistent Laplacians derived from algebraic topology. In contrast to graph Laplacians, persistent Laplacians enable multiscale analysis through filtration and incorporate higher-order simplicial complexes to capture higher-order interactions in the data. We evaluate and validate the performance of PLPCA using benchmark microarray datasets that involve normal tissue samples and four different cancer tissues. Our extensive studies demonstrate that PLPCA outperforms all other state-of-the-art models for classification tasks after dimensionality reduction.

math.AT

Virtual screening of DrugBank database for hERG blockers using topological Laplacian-assisted AI models

The human {\it ether-a-go-go} (hERG) potassium channel (K$_\text{v}11.1$) plays a critical role in mediating cardiac action potential. The blockade of this ion channel can potentially lead fatal disorder and/or long QT syndrome. Many drugs have been withdrawn because of their serious hERG-cardiotoxicity. It is crucial to assess the hERG blockade activity in the early stage of drug discovery. We are particularly interested in the hERG-cardiotoxicity of compounds collected in the DrugBank database considering that many DrugBank compounds have been approved for therapeutic treatments or have high potential to become drugs. Machine learning-based in silico tools offer a rapid and economical platform to virtually screen DrugBank compounds. We design accurate and robust classifiers for blockers/non-blockers and then build regressors to quantitatively analyze the binding potency of the DrugBank compounds on the hERG channel. Molecular sequences are embedded with two natural language processing (NPL) methods, namely, autoencoder and transformer. Complementary three-dimensional (3D) molecular structures are embedded with two advanced mathematical approaches, i.e., topological Laplacians and algebraic graphs. With our state-of-the-art tools, we reveal that 227 out of the 8641 DrugBank compounds are potential hERG blockers, suggesting serious drug safety problems. Our predictions provide guidance for the further experimental interrogation of DrugBank compounds' hERG-cardiotoxicity .

q-bio.BM

Prediction and mitigation of mutation threats to COVID-19 vaccines and antibody therapies

Antibody therapeutics and vaccines are among our last resort to end the raging COVID-19 pandemic. They, however, are prone to over 5,000 mutations on the spike (S) protein uncovered by a Mutation Tracker based on over 200,000 genome isolates. It is imperative to understand how mutations would impact vaccines and antibodies in the development. In this work, we study the mechanism, frequency, and ratio of mutations on the S protein. Additionally, we use 56 antibody structures and analyze their 2D and 3D characteristics. Moreover, we predict the mutation-induced binding free energy (BFE) changes for the complexes of S protein and antibodies or ACE2. By integrating genetics, biophysics, deep learning, and algebraic topology, we reveal that most of 462 mutations on the receptor-binding domain (RBD) will weaken the binding of S protein and antibodies and disrupt the efficacy and reliability of antibody therapies and vaccines. A list of 31 vaccine escape mutants is identified, while many other disruptive mutations are detailed as well. We also unveil that about 65\% existing RBD mutations, including those variants recently found in the United Kingdom (UK) and South Africa, are binding-strengthen mutations, resulting in more infectious COVID-19 variants. We discover the disparity between the extreme values of RBD mutation-induced BFE strengthening and weakening of the bindings with antibodies and ACE2, suggesting that SARS-CoV-2 is at an advanced stage of evolution for human infection, while the human immune system is able to produce optimized antibodies. This discovery implies the vulnerability of current vaccines and antibody drugs to new mutations. Our predictions were validated by comparison with more than 1,400 deep mutations on the S protein RBD. Our results show the urgent need to develop new mutation-resistant vaccines and antibodies and to prepare for seasonal vaccinations.

q-bio.BM

Representability of algebraic topology for biomolecules in machine learning based scoring and virtual screening

This work introduces a number of algebraic topology approaches, such as multicomponent persistent homology, multi-level persistent homology and electrostatic persistence for the representation, characterization, and description of small molecules and biomolecular complexes. Multicomponent persistent homology retains critical chemical and biological information during the topological simplification of biomolecular geometric complexity. Multi-level persistent homology enables a tailored topological description of inter- and/or intra-molecular interactions of interest. Electrostatic persistence incorporates partial charge information into topological invariants. These topological methods are paired with Wasserstein distance to characterize similarities between molecules and are further integrated with a variety of machine learning algorithms, including k-nearest neighbors, ensemble of trees, and deep convolutional neural networks, to manifest their descriptive and predictive powers for chemical and biological problems. Extensive numerical experiments involving more than 4,000 protein-ligand complexes from the PDBBind database and near 100,000 ligands and decoys in the DUD database are performed to test respectively the scoring power and the virtual screening power of the proposed topological approaches. It is demonstrated that the present approaches outperform the modern machine learning based methods in protein-ligand binding affinity predictions and ligand-decoy discrimination.

q-bio.QM

A Review of Mathematical Modeling, Simulation and Analysis of Membrane Channel Charge Transport

The molecular mechanism of ion channel gating and substrate modulation is elusive for many voltage gated ion channels, such as eukaryotic sodium ones. The understanding of channel functions is a pressing issue in molecular biophysics and biology. Mathematical modeling, computation and analysis of membrane channel charge transport have become an emergent field and give rise to significant contributions to our understanding of ion channel gating and function. This review summarizes recent progresses and outlines remaining challenges in mathematical modeling, simulation and analysis of ion channel charge transport. One of our focuses is the Poisson-Nernst-Planck (PNP) model and its generalizations. Specifically, the basic framework of the PNP system and some of its extensions, including size effects, ion-water interactions, coupling with density functional theory and relation to fluid flow models. A reduced theory, the Poisson- Boltzmann-Nernst-Planck (PBNP) model, and a differential geometry based ion transport model are also discussed. For proton channel, a multiscale and multiphysics Poisson-Boltzmann-Kohn-Sham (PBKS) model is presented. We show that all of these ion channel models can be cast into a unified variational multiscale framework with a macroscopic continuum domain of the solvent and a microscopic discrete domain of the solute. The main strategy is to construct a total energy functional of a charge transport system to encompass the polar and nonpolar free energies of solvation and chemical potential related energies. Current computational algorithms and tools for numerical simulations and results from mathematical analysis of ion channel systems are also surveyed.

q-bio.BM

Accurate, robust and reliable calculations of Poisson-Boltzmann solvation energies

Developing accurate solvers for the Poisson Boltzmann (PB) model is the first step to make the PB model suitable for implicit solvent simulation. Reducing the grid size influence on the performance of the solver benefits to increasing the speed of solver and providing accurate electrostatics analysis for solvated molecules. In this work, we explore the accurate coarse grid PB solver based on the Green's function treatment of the singular charges, matched interface and boundary (MIB) method for treating the geometric singularities, and posterior electrostatic potential field extension for calculating the reaction field energy. We made our previous PB software, MIBPB, robust and provides almost grid size independent reaction field energy calculation. Large amount of the numerical tests verify the grid size independence merit of the MIBPB software. The advantage of MIBPB software directly make the acceleration of the PB solver from the numerical algorithm instead of utilization of advanced computer architectures. Furthermore, the presented MIBPB software is provided as a free online sever.

math.NA

Automatic parametrization of implicit solvent models for the blind prediction of solvation free energies

In this work, a systematic protocol is proposed to automatically parametrize implicit solvent models with polar and nonpolar components. The proposed protocol utilizes the classical Poisson model or the Kohn-Sham density functional theory (KSDFT) based polarizable Poisson model for modeling polar solvation free energies. For the nonpolar component, either the standard model of surface area, molecular volume, and van der Waals interactions, or a model with atomic surface areas and molecular volume is employed. Based on the assumption that similar molecules have similar parametrizations, we develop scoring and ranking algorithms to classify solute molecules. Four sets of radius parameters are combined with four sets of charge force fields to arrive at a total of 16 different parametrizations for the Poisson model. A large database with 668 experimental data is utilized to validate the proposed protocol. The lowest leave-one-out root mean square (RMS) error for the database is 1.33k cal/mol. Additionally, five subsets of the database, i.e., SAMPL0-SAMPL4, are employed to further demonstrate that the proposed protocol offers some of the best solvation predictions. The optimal RMS errors are 0.93, 2.82, 1.90, 0.78, and 1.03 kcal/mol, respectively for SAMPL0, SAMPL1, SAMPL2, SAMPL3, and SAMPL4 test sets. These results are some of the best, to our best knowledge.

physics.chem-ph

Parameter optimization in differential geometry based solvation models

Differential geometry (DG) based solvation models are a new class of variational implicit solvent approaches that are able to avoid unphysical solvent-solute boundary definitions and associated geometric singularities, and dynamically couple polar and nonpolar interactions in a self-consistent framework. Our earlier study indicates that DG based nonpolar solvation model outperforms other methods in nonpolar solvation energy predictions. However, the DG based full solvation model has not shown its superiority in solvation analysis, due to its difficulty in parametrization, which must ensure the stability of the solution of strongly coupled nonlinear Laplace-Beltrami and Poisson-Boltzmann equations. In this work, we introduce new parameter learning algorithms based on perturbation and convex optimization theories to stabilize the numerical solution and thus achieve an optimal parametrization of the DG based solvation models. An interesting feature of the present DG based solvation model is that it provides accurate solvation free energy predictions for both polar and nonploar molecules in a unified formulation. Extensive numerical experiment demonstrates that the present DG based solvation model delivers some of the most accurate predictions of the solvation free energies for a large number of molecules.

math.NA

Second order Method for Solving 3D Elasticity Equations with Complex and Sharp Interfaces

Elastic materials are ubiquitous in nature and indispensable components in man-made devices and equipments. When a device or equipment involves composite or multiple elastic materials, elasticity interface problems come into play. The solution of three dimensional (3D) elasticity interface problems is significantly more difficult than that of elliptic counterparts due to the coupled vector components and cross derivatives in the governing elasticity equation. This work introduces the matched interface and boundary (MIB) method for solving 3D elasticity interface problems. The proposed MIB method utilizes fictitious values on irregular grid points near the material interface to replace function values in the discretization so that the elasticity equation can be discretized using the standard finite difference schemes as if there were no material interface. The interface jump conditions are rigorously enforced on the intersecting points between the interface and the mesh lines. Such an enforcement determines the fictitious values. A number of new technique are developed to construct efficient MIB schemes for dealing with cross derivative in coupled governing equations. The proposed method is extensively validated over both weak and strong discontinuity of the solution, both piecewise constant and position-dependent material parameters, both smooth and nonsmooth interface geometries, and both small and large contrasts in the Poisson's ratio and shear modulus across the interface. Numerical experiments indicate that the present MIB method is of second order convergence in both $L_\infty$ and $L_2$ error norms.

math.NA

Weak Galerkin Methods for Second Order Elliptic Interface Problems

Weak Galerkin methods refer to general finite element methods for PDEs in which differential operators are approximated by their weak forms as distributions. Such weak forms give rise to desirable flexibilities in enforcing boundary and interface conditions. A weak Galerkin finite element method (WG-FEM) is developed in this paper for solving elliptic partial differential equations (PDEs) with discontinuous coefficients and interfaces. The paper also presents many numerical tests for validating the WG-FEM for solving second order elliptic interface problems. For such interface problems, the solution possesses a certain singularity due to the nonsmoothness of the interface. A challenge in research is to design high order numerical methods that work well for problems with low regularity in the solution. The best known numerical scheme in the literature is of order one for the solution itself in $L_\infty$ norm. It is demonstrated that the WG-FEM of lowest order is capable of delivering numerical approximations that are of order 1.75 in the usual $L_\infty$ norm for $C^1$ or Lipschitz continuous interfaces associated with a $C^1$ or $H^2$ continuous solutions. Theoretically, it is proved that high order of numerical schemes can be designed by using the WG-FEM with polynomials of high order on each element.

math.NA