SearcharxivSearch

arXiv subjects

Washington Mio

Publications and source records attributed to Washington Mio.

18 recordsLinked to original sources

The Observable Wasserstein Distance

We introduce the observable Wasserstein distance, a framework for deriving lower bounds on the Wasserstein distance between probability measures on Polish metric spaces, designed to bypass the computational intractability of exact optimal transport in large-scale, non-Euclidean datasets. Analogous to the sliced Wasserstein distance in $\mathbb{R}^d$, our approach projects measures onto the real line via 1-Lipschitz observables and computes the Wasserstein distances between the resulting pushforward distributions. We define a hierarchy of pseudo-metrics by restricting observables to a nested chain of subspaces. A central theoretical contribution is an injectivity result linking the metric covering dimension of the support of a measure to the specific order in the hierarchy that guarantees unique recovery. This serves as a metric-space analogue to the Cram\'{e}r-Wold Device for Euclidean distributions. We demonstrate that this hierarchy offers a tunable trade-off between sharpness as a lower bound on the Wasserstein distance and computational efficiency. We also present a discrete computational model for finite grids and numerical experiments validating the efficacy and utility of these approximations.

math.MG

Observable Covariance and Principal Observable Analysis for Data on Metric Spaces

Datasets consisting of objects such as shapes, networks, images, or signals overlaid on such geometric objects permeate data science. Such datasets are often equipped with metrics that quantify the similarity or divergence between any pair of elements turning them into metric spaces $(X,d)$, or a metric measure space $(X,d,\mu)$ if data density is also accounted for through a probability measure $\mu$. This paper develops a Lipschitz geometry approach to analysis of metric measure spaces based on metric observables; that is, 1-Lipschitz scalar fields $f \colon X \to \mathbb{R}$ that provide reductions of $(X,d,\mu)$ to $\mathbb{R}$ through the projected measure $f_\sharp (\mu)$. Collectively, metric observables capture a wealth of information about the shape of $(X,d,\mu)$ at all spatial scales. In particular, we can define stable statistics such as the observable mean and observable covariance operators $M_\mu$ and $\Sigma_\mu$, respectively. Through a maximization of variance principle, analogous to principal component analysis, $\Sigma_\mu$ leads to an approach to vectorization, dimension reduction, and visualization of metric measure data that we term principal observable analysis. The method also yields basis functions for representation of signals on $X$ in the observable domain.

math.ST

Robust Representation and Estimation of Barycenters and Modes of Probability Measures on Metric Spaces

This paper is concerned with the problem of defining and estimating statistics for distributions on spaces such as Riemannian manifolds and more general metric spaces. The challenge comes, in part, from the fact that statistics such as means and modes may be unstable: for example, a small perturbation to a distribution can lead to a large change in Fr\'echet means on spaces as simple as a circle. We address this issue by introducing a new merge tree representation of barycenters called the barycentric merge tree (BMT), which takes the form of a measured metric graph and summarizes features of the distribution in a multiscale manner. Modes are treated as special cases of barycenters through diffusion distances. In contrast to the properties of classical means and modes, we prove that BMTs are stable -- this is quantified as a Lipschitz estimate involving optimal transport metrics. This stability allows us to derive a consistency result for approximating BMTs from empirical measures, with explicit convergence rates. We also give a provably accurate method for discretely approximating the BMT construction and use this to provide numerical examples for distributions on spheres and shape spaces.

math.ST

Stability and Approximations for Decorated Reeb Spaces

Given a map $f:X \to M$ from a topological space $X$ to a metric space $M$, a decorated Reeb space consists of the Reeb space, together with an attribution function whose values recover geometric information lost during the construction of the Reeb space. For example, when $M=\mathbb{R}$ is the real line, the Reeb space is the well-known Reeb graph, and the attributions may consist of persistence diagrams summarizing the level set topology of $f$. In this paper, we introduce decorated Reeb spaces in various flavors and prove that our constructions are Gromov-Hausdorff stable. We also provide results on approximating decorated Reeb spaces from finite samples and leverage these to develop a computational framework for applying these constructions to point cloud data.

math.MG

On Metrics for Analysis of Functional Data on Geometric Domains

This paper employs techniques from metric geometry and optimal transport theory to address questions related to the analysis of functional data on metric or metric-measure spaces, which we refer to as fields. Formally, fields are viewed as 1-Lipschitz mappings between Polish metric spaces with the domain possibly equipped with a Borel probability measure. We introduce field analogues of the Gromov-Hausdorff, Gromov-Prokhorov, and Gromov-Wasserstein distances, investigate their main properties and provide a characterization of the Gromov-Hausdorff distance in terms of isometric embeddings in a Urysohn universal field. Adapting the notion of distance matrices to fields, we formulate a discrete model, obtain an empirical estimation result that provides a theoretical basis for its use in functional data analysis, and prove a field analogue of Gromov's Reconstruction Theorem. We also investigate field versions of the Vietoris-Rips and neighborhood (or offset) filtrations and prove that they are stable with respect to appropriate metrics.

math.MG

Topologically Attributed Graphs for Shape Discrimination

In this paper we introduce a novel family of attributed graphs for the purpose of shape discrimination. Our graphs typically arise from variations on the Mapper graph construction, which is an approximation of the Reeb graph for point cloud data. Our attributions enrich these constructions with (persistent) homology in ways that are provably stable, thereby recording extra topological information that is typically lost in these graph constructions. We provide experiments which illustrate the use of these invariants for shape representation and classification. In particular, we obtain competitive shape classification results when using our topologically attributed graphs as inputs to a simple graph neural network classifier.

math.AT

Convergence of Leray Cosheaves for Decorated Mapper Graphs

We introduce decorated mapper graphs as a generalization of mapper graphs capable of capturing more topological information of a data set. A decorated mapper graph can be viewed as a discrete approximation of the cellular Leray cosheaf over the Reeb graph. We establish a theoretical foundation for this construction by showing that the cellular Leray cosheaf with respect to a sequence of covers converges to the actual Leray cosheaf as the resolution of the covers goes to zero.

math.AT

Universal Mappings and Analysis of Functional Data on Geometric Domains

This paper addresses problems in functional metric geometry that arise in the study of data such as signals recorded on geometric domains or on the nodes of weighted networks. Datasets comprising such objects arise in many domains of scientific and practical interest. For example, $f$ could represent a functional magnetic resonance image, or the nodes of a social network labeled with attributes or preferences, where the underlying metric structure is given by the shortest path distance, commute distance, or diffusion distance. Formally, these may be viewed as functions defined on metric spaces, sometimes equipped with additional structure such as a probability measure, in which case the domain is referred to as a metric-measure space, or simply $mm$-space. Our primary goal is threefold: (i) to develop metrics that allow us to model and quantify variation in functional data, possibly with distinct domains; (ii) to investigate principled empirical estimations of these metrics; (iii) to construct a universal function that ``contains'' all functions whose domains and ranges are Polish (separable and complete metric) spaces, assuming Lipschitz regularity. The latter is much in the spirit of constructing universal spaces for structural data (metric spaces) whose investigation dates back to the early 20th century and are of classical interest in metric geometry.

math.MG

Detecting Carbon Nanotube Orientation with Topological Data Analysis of Scanning Electron Micrographs

As the aerospace industry becomes increasingly demanding for stronger lightweight materials, the ultra-strong carbon nanotube (CNT) composites with highly aligned CNT network structures could be the answer. In this work, a novel methodology applying topological data analysis (TDA) to the scanning electron microscope (SEM) images was developed to detect CNT orientation. The CNT bundle extensions in certain directions were summarized algebraically and expressed as visible barcodes. The barcodes were then calculated and converted into the total spread function $V(X,θ)$, from which the alignment fraction and the preferred direction could be determined. For validation purposes, the random CNT sheets were mechanically stretched at various strain ratios ranging from $0-40\%$, and quantitative TDA analysis was conducted based on the SEM images taken at random positions. The results showed high consistency ($R^2=0.975$) compared to the Herman's orientation factors derived from the polarized Raman spectroscopy and wide-angle X-ray scattering analysis. Additionally, the TDA method presented great robustness with varying SEM acceleration voltages and magnifications, which might alter the scope in alignment detection. With potential applications in nanofiber systems, this study offers a rapid and simple way to quantify CNT alignment, which plays a crucial role in transferring the CNT properties into engineering products.

cond-mat.mtrl-sci

Decorated Merge Trees for Persistent Topology

This paper introduces decorated merge trees (DMTs) as a novel invariant for persistent spaces. DMTs combine both $π_0$ and $H_n$ information into a single data structure that distinguishes filtrations that merge trees and persistent homology cannot distinguish alone. Three variants on DMTs, which emphasize category theory, representation theory and persistence barcodes, respectively, offer different advantages in terms of theory and computation. Two notions of distance -- an interleaving distance and bottleneck distance -- for DMTs are defined and a hierarchy of stability results that both refine and generalize existing stability results is proved here. To overcome some of the computational complexity inherent in these distances, we provide a novel use of Gromov-Wasserstein couplings to compute optimal merge tree alignments for a combinatorial version of our interleaving distance which can be tractably estimated. We introduce computational frameworks for generating, visualizing and comparing decorated merge trees derived from synthetic and real data. Example applications include comparison of point clouds, interpretation of persistent homology of sliding window embeddings of time series, visualization of topological features in segmented brain tumor images and topology-driven graph alignment.

math.AT

Correspondence Modules and Persistence Sheaves: A Unifying Perspective on One-Parameter Persistent Homology

We develop a unifying framework for the treatment of various persistent homology architectures using the notion of correspondence modules. In this formulation, morphisms between vector spaces are given by partial linear relations, as opposed to linear mappings. In the one-dimensional case, among other things, this allows us to: (i) treat persistence modules and zigzag modules as algebraic objects of the same type; (ii) give a categorical formulation of zigzag structures over a continuous parameter; and (iii) construct barcodes associated with spaces and mappings that are richer in geometric information. A structural analysis of one-parameter persistence is carried out at the level of sections of correspondence modules that yield sheaf-like structures, termed persistence sheaves. Under some tameness hypotheses, we prove interval decomposition theorems for persistence sheaves and correspondence modules, as well as an isometry theorem for persistence diagrams obtained from interval decompositions. Applications include: (a) a Mayer-Vietoris sequence that relates the persistent homology of sublevelset filtrations and superlevelset filtrations to the levelset homology module of a real-valued function and (b) the construction of slices of 2-parameter persistence modules along negatively sloped lines.

math.AT

Towards Quantifying Intrinsic Generalization of Deep ReLU Networks

Understanding the underlying mechanisms that enable the empirical successes of deep neural networks is essential for further improving their performance and explaining such networks. Towards this goal, a specific question is how to explain the "surprising" behavior of the same over-parametrized deep neural networks that can generalize well on real datasets and at the same time "memorize" training samples when the labels are randomized. In this paper, we demonstrate that deep ReLU networks generalize from training samples to new points via piece-wise linear interpolation. We provide a quantified analysis on the generalization ability of a deep ReLU network: Given a fixed point $\mathbf{x}$ and a fixed direction in the input space $\mathcal{S}$, there is always a segment such that any point on the segment will be classified the same as the fixed point $\mathbf{x}$. We call this segment the $generalization \ interval$. We show that the generalization intervals of a ReLU network behave similarly along pairwise directions between samples of the same label in both real and random cases on the MNIST and CIFAR-10 datasets. This result suggests that the same interpolation mechanism is used in both cases. Additionally, for datasets using real labels, such networks provide a good approximation of the underlying manifold in the data, where the changes are much smaller along tangent directions than along normal directions. On the other hand, however, for datasets with random labels, generalization intervals along mid-lines of triangles with the same label are much smaller than those on the datasets with real labels, suggesting different behaviors along other directions. Our systematic experiments demonstrate for the first time that such deep neural networks generalize through the same interpolation and explain the differences between their performance on datasets with real and random labels.

cs.LG

A Topological Study of Functional Data and Fréchet Functions of Metric Measure Spaces

We study the persistent homology of both functional data on compact topological spaces and structural data presented as compact metric measure spaces. One of our goals is to define persistent homology so as to capture primarily properties of the shape of a signal, eliminating otherwise highly persistent homology classes that may exist simply because of the nature of the domain on which the signal is defined. We investigate the stability of these invariants using metrics that downplay regions where signals are weak. The distance between two signals is small if they exhibit high similarity in regions where they are strong, regardless of the nature of their full domains, in particular allowing different homotopy types. Consistency and estimation of persistent homology of metric measure spaces from data are studied within this framework. We also apply the methodology to the construction of multi-scale topological descriptors for data on compact Riemannian manifolds via metric relaxations derived from the heat kernel.

math.AT

Probing the Geometry of Data with Diffusion Fréchet Functions

Many complex ecosystems, such as those formed by multiple microbial taxa, involve intricate interactions amongst various sub-communities. The most basic relationships are frequently modeled as co-occurrence networks in which the nodes represent the various players in the community and the weighted edges encode levels of interaction. In this setting, the composition of a community may be viewed as a probability distribution on the nodes of the network. This paper develops methods for modeling the organization of such data, as well as their Euclidean counterparts, across spatial scales. Using the notion of diffusion distance, we introduce diffusion Frechet functions and diffusion Frechet vectors associated with probability distributions on Euclidean space and the vertex set of a weighted network, respectively. We prove that these functional statistics are stable with respect to the Wasserstein distance between probability measures, thus yielding robust descriptors of their shapes. We apply the methodology to investigate bacterial communities in the human gut, seeking to characterize divergence from intestinal homeostasis in patients with Clostridium difficile infection (CDI) and the effects of fecal microbiota transplantation, a treatment used in CDI patients that has proven to be significantly more effective than traditional treatment with antibiotics. The proposed method proves useful in deriving a biomarker that might help elucidate the mechanisms that drive these processes.

stat.ML

The Shape of Data and Probability Measures

We introduce the notion of multiscale covariance tensor fields (CTF) associated with Euclidean random variables as a gateway to the shape of their distributions. Multiscale CTFs quantify variation of the data about every point in the data landscape at all spatial scales, unlike the usual covariance tensor that only quantifies global variation about the mean. Empirical forms of localized covariance previously have been used in data analysis and visualization, but we develop a framework for the systematic treatment of theoretical questions and computational models based on localized covariance. We prove strong stability theorems with respect to the Wasserstein distance between probability measures, obtain consistency results, as well as estimates for the rate of convergence of empirical CTFs. These results ensure that CTFs are robust to sampling, noise and outliers. We provide numerous illustrations of how CTFs let us extract shape from data and also apply CTFs to manifold clustering, the problem of categorizing data points according to their noisy membership in a collection of possibly intersecting, smooth submanifolds of Euclidean space. We prove that the proposed manifold clustering method is stable and carry out several experiments to validate the method.

stat.ML

The Phylogenetic LASSO and the Microbiome

Scientific investigations that incorporate next generation sequencing involve analyses of high-dimensional data where the need to organize, collate and interpret the outcomes are pressingly important. Currently, data can be collected at the microbiome level leading to the possibility of personalized medicine whereby treatments can be tailored at this scale. In this paper, we lay down a statistical framework for this type of analysis with a view toward synthesis of products tailored to individual patients. Although the paper applies the technique to data for a particular infectious disease, the methodology is sufficiently rich to be expanded to other problems in medicine, especially those in which coincident `-omics' covariates and clinical responses are simultaneously captured.

stat.ML

The quadratic form E_8 and exotic homology manifolds

An explicit (-1)^n-quadratic form over Z[Z^{2n}] representing the surgery problem E_8 x T^{2n} is obtained, for use in the Bryant-Ferry-Mio-Weinberger construction of 2n-dimensional exotic homology manifolds.

math.GT

Topology of homology manifolds

We construct examples of nonresolvable generalized $n$-manifolds, $n\geq 6$, with arbitrary resolution obstruction, homotopy equivalent to any simply connected, closed $n$-manifold. We further investigate the structure of generalized manifolds and present a program for understanding their topology.

math.GT