Searcharxiv⌕ Search

arXiv subjects

Tom Needham

Publications and source records attributed to Tom Needham.

At least 37 records · Page 2Linked to original sources

Geometry of the Space of Partitioned Networks: A Unified Theoretical and Computational Framework

Interactions and relations between objects may be pairwise or higher-order in nature, and so network-valued data are ubiquitous in the real world. The "space of networks", however, has a complex structure that cannot be adequately described using conventional statistical tools. We introduce a measure-theoretic formalism for modeling generalized network structures such as graphs, hypergraphs, or graphs whose nodes come with a partition into categorical classes. We then propose a metric that extends the Gromov-Wasserstein distance between graphs and the co-optimal transport distance between hypergraphs. We characterize the geometry of this space, thereby providing a unified theoretical treatment of generalized networks that encompasses the cases of pairwise, as well as higher-order, relations. In particular, we show that our metric is an Alexandrov space of non-negative curvature, and leverage this structure to define gradients for certain functionals commonly arising in geometric data analysis tasks. We extend our analysis to the setting where vertices have additional label information, and derive efficient computational schemes to use in practice. Equipped with these theoretical and computational tools, we demonstrate the utility of our framework in a suite of applications, including hypergraph alignment, clustering and dictionary learning from ensemble data, multi-omics alignment, as well as multiscale network alignment.

math.MG↗

Stability of Hypergraph Invariants and Transformations

Graphs are fundamental tools for modeling pairwise interactions in complex systems. However, many real-world systems involve multi-way interactions that cannot be fully captured by standard graphs. Hypergraphs, which generalize graphs by allowing edges to connect any number of vertices, offer a more expressive framework. In this paper, we introduce a new metric on the space of hypergraphs, inspired by the Gromov-Hausdorff distance for metric spaces. We establish Lipschitz properties of common hypergraph transformations, which send hypergraphs to graphs, including a novel graphification method with ties to single linkage hierarchical clustering. Additionally, we derive lower bounds for the hypergraph distance via invariants coming from basic summary statistics and from topological data analysis techniques. Finally, we explore stability properties of cost functions in the context of optimal transport. Our results in this direction consider Lipschitzness of the Hausdorff map and conservation of the non-negative cross curvature property under limits of cost functions.

math.MG↗

Fused Gromov-Wasserstein Variance Decomposition with Linear Optimal Transport

Wasserstein distances form a family of metrics on spaces of probability measures that have recently seen many applications. However, statistical analysis in these spaces is complex due to the nonlinearity of Wasserstein spaces. One potential solution to this problem is Linear Optimal Transport (LOT). This method allows one to find a Euclidean embedding, called LOT embedding, of measures in some Wasserstein spaces, but some information is lost in this embedding. So, to understand whether statistical analysis relying on LOT embeddings can make valid inferences about original data, it is helpful to quantify how well these embeddings describe that data. To answer this question, we present a decomposition of the Fréchet variance of a set of measures in the 2-Wasserstein space, which allows one to compute the percentage of variance explained by LOT embeddings of those measures. We then extend this decomposition to the Fused Gromov-Wasserstein setting. We also present several experiments that explore the relationship between the dimension of the LOT embedding, the percentage of variance explained by the embedding, and the classification accuracy of machine learning classifiers built on the embedded data. We use the MNIST handwritten digits dataset, IMDB-50000 dataset, and Diffusion Tensor MRI images for these experiments. Our results illustrate the effectiveness of low dimensional LOT embeddings in terms of the percentage of variance explained and the classification accuracy of models built on the embedded data.

stat.ME↗

Comparison Results for Gromov-Wasserstein and Gromov-Monge Distances

Inspired by the Kantorovich formulation of optimal transport distance between probability measures on a metric space, Gromov-Wasserstein (GW) distances comprise a family of metrics on the space of isomorphism classes of metric measure spaces. In previous work, the authors introduced a variant of this construction which was inspired by the original Monge formulation of optimal transport; elements of the resulting family are referred to Gromov-Monge (GM) distances. These GM distances, and related ideas, have since become a subject of interest from both theoretical and applications-oriented perspectives. In this note, we establish several theoretical properties of GM distances, focusing on comparisons between GM and GW distances. In particular, we show that GM and GW distances are equal for non-atomic metric measure spaces. We also consider variants of GM distance, such as a Monge version of Sturm's $L_p$-transportion distance, and give precise comparisons to GW distance. Finally, we establish bi-Hölder equivalence between GM distance and an isometry-invariant Monge optimal transport distance between Euclidean metric measure spaces that has been utilized in shape and image analysis applications.

math.MG↗

Geometric Approaches to Matrix Normalization and Graph Balancing

Normal matrices, or matrices which commute with their adjoints, are of fundamental importance in pure and applied mathematics. In this paper, we study a natural functional on the space of square complex matrices whose global minimizers are normal matrices. We show that this functional, which we refer to as the non-normal energy, has incredibly well-behaved gradient descent dynamics: despite it being non-convex, we show that the only critical points of the non-normal energy are the normal matrices, and that its gradient descent trajectories fix matrix spectra and preserve the subset of real matrices. We also show that, even when restricted to the subset of unit Frobenius norm matrices, the gradient flow of the non-normal energy retains many of these useful properties. This is applied to prove that low-dimensional homotopy groups of spaces of unit norm normal matrices vanish; for example, we show that the space of $d \times d$ complex unit norm normal matrices is simply connected for all $d \geq 2$. Finally, we consider the related problem of balancing a weighted directed graph -- that is, readjusting its edge weights so that the weighted in-degree and out-degree is the same at each node. We adapt the non-normal energy to define another natural functional whose global minima are balanced graphs and show that gradient descent of this functional always converges to a balanced graph, while preserving graph spectra and realness of the weights. Our results were inspired by concepts from symplectic geometry and Geometric Invariant Theory, but we mostly avoid invoking this machinery and our proofs are generally self-contained.

math.DG↗

Stability and Approximations for Decorated Reeb Spaces

Given a map $f:X \to M$ from a topological space $X$ to a metric space $M$, a decorated Reeb space consists of the Reeb space, together with an attribution function whose values recover geometric information lost during the construction of the Reeb space. For example, when $M=\mathbb{R}$ is the real line, the Reeb space is the well-known Reeb graph, and the attributions may consist of persistence diagrams summarizing the level set topology of $f$. In this paper, we introduce decorated Reeb spaces in various flavors and prove that our constructions are Gromov-Hausdorff stable. We also provide results on approximating decorated Reeb spaces from finite samples and leverage these to develop a computational framework for applying these constructions to point cloud data.

math.MG↗

On the existence of Parseval frames for vector bundles

Frames in finite-dimensional vector spaces are spanning sets of vectors which provide redundant representations of signals. The Parseval frames are particularly useful and important, since they provide a simple reconstruction scheme and are maximally robust against certain types of noise. In this paper we describe a theory of frames on arbitrary vector bundles -- this is the natural setting for signals which are realized as parameterized families of vectors rather than as single vectors -- and discuss the existence of Parseval frames in this setting. Our approach is phrased in the language of $G$-bundles, which allows us to use many tools from classical algebraic topology. In particular, we show that orientable vector bundles always admit Parseval frames of sufficiently large size and provide an upper bound on the necessary size. We also give sufficient conditions for the existence of Parseval frames of smaller size for tangent bundles of several families of manifolds, and provide some numerical evidence that Parseval frames on vector bundles share the desirable reconstruction properties of classical Parseval frames.

math.DG↗

A Wasserstein-type Distance for Gaussian Mixtures on Vector Bundles with Applications to Shape Analysis

This paper uses sample data to study the problem of comparing populations on finite-dimensional parallelizable Riemannian manifolds and more general trivial vector bundles. Utilizing triviality, our framework represents populations as mixtures of Gaussians on vector bundles and estimates the population parameters using a mode-based clustering algorithm. We derive a Wasserstein-type metric between Gaussian mixtures, adapted to the manifold geometry, in order to compare estimated distributions. Our contributions include an identifiability result for Gaussian mixtures on manifold domains and a convenient characterization of optimal couplings of Gaussian mixtures under the derived metric. We demonstrate these tools on some example domains, including the pre-shape space of planar closed curves, with applications to the shape space of triangles and populations of nanoparticles. In the nanoparticle application, we consider a sequence of populations of particle shapes arising from a manufacturing process, and utilize the Wasserstein-type distance to perform change-point detection.

stat.ME↗

Interleaving Distances, Monoidal Actions and 2-Categories

Interleaving distances are used widely in Topological Data Analysis (TDA) as a tool for comparing topological signatures of datasets. The theory of interleaving distances has been extended through various category-theoretic constructions, enabling its usage beyond standard constructions of TDA, while clarifying certain observed stability phenomena by unifying them under a common framework. Inspired by metrics used in the field of statistical shape analysis, which are based on minimizing energy functions over group actions, we define three new types of increasingly general interleaving distances. Our constructions use ideas from the theories of monoidal actions and 2-categories. We show that these distances naturally extend the category with a flow framework of de Silva, Munch and Stefanou and the locally persistent category framework of Scoccola, and we provide a general stability result. Along the way, we give examples of distances that fit into our framework which connect to ideas from differential geometry, geometric shape analysis, statistical TDA and multiparameter persistent homology.

math.AT↗

Comparing Morse Complexes Using Optimal Transport: An Experimental Study

Morse complexes and Morse-Smale complexes are topological descriptors popular in topology-based visualization. Comparing these complexes plays an important role in their applications in feature correspondences, feature tracking, symmetry detection, and uncertainty visualization. Leveraging recent advances in optimal transport, we apply a class of optimal transport distances to the comparative analysis of Morse complexes. Contrasting with existing comparative measures, such distances are easy and efficient to compute, and naturally provide structural matching between Morse complexes. We perform an experimental study involving scientific simulation datasets and discuss the effectiveness of these distances as comparative measures for Morse complexes. We also provide an initial guideline for choosing the optimal transport distances under various data assumptions.

cs.CG↗

Topologically Attributed Graphs for Shape Discrimination

In this paper we introduce a novel family of attributed graphs for the purpose of shape discrimination. Our graphs typically arise from variations on the Mapper graph construction, which is an approximation of the Reeb graph for point cloud data. Our attributions enrich these constructions with (persistent) homology in ways that are provably stable, thereby recording extra topological information that is typically lost in these graph constructions. We provide experiments which illustrate the use of these invariants for shape representation and classification. In particular, we obtain competitive shape classification results when using our topologically attributed graphs as inputs to a simple graph neural network classifier.

math.AT↗

Convergence of Leray Cosheaves for Decorated Mapper Graphs

We introduce decorated mapper graphs as a generalization of mapper graphs capable of capturing more topological information of a data set. A decorated mapper graph can be viewed as a discrete approximation of the cellular Leray cosheaf over the Reeb graph. We establish a theoretical foundation for this construction by showing that the cellular Leray cosheaf with respect to a sequence of covers converges to the actual Leray cosheaf as the resolution of the covers goes to zero.

math.AT↗

Hypergraph Co-Optimal Transport: Metric and Categorical Properties

Hypergraphs capture multi-way relationships in data, and they have consequently seen a number of applications in higher-order network analysis, computer vision, geometry processing, and machine learning. In this paper, we develop theoretical foundations for studying the space of hypergraphs using ingredients from optimal transport. By enriching a hypergraph with probability measures on its nodes and hyperedges, as well as relational information capturing local and global structures, we obtain a general and robust framework for studying the collection of all hypergraphs. First, we introduce a hypergraph distance based on the co-optimal transport framework of Redko et al. and study its theoretical properties. Second, we formalize common methods for transforming a hypergraph into a graph as maps between the space of hypergraphs and the space of graphs, and study their functorial properties and Lipschitz bounds. Finally, we demonstrate the versatility of our Hypergraph Co-Optimal Transport (HyperCOT) framework through various examples.

math.MG↗

Statistical Shape Analysis of Shape Graphs with Applications to Retinal Blood-Vessel Networks

This paper provides theoretical and computational developments in statistical shape analysis of shape graphs, and demonstrates them using analysis of complex data from retinal blood-vessel (RBV) networks. The shape graphs are represented by a set of nodes and edges (planar articulated curves) connecting some of these nodes. The goals are to utilize shapes of edges and connectivities and locations of nodes to: (1) characterize full shapes, (2) quantify shape differences, and (3) model statistical variability. We develop a mathematical representation, elastic Riemannian shape metrics, and associated tools for such statistical analysis. Specifically, we derive tools for shape graph registration, geodesics, summaries, and shape modeling. Geodesics are convenient for visualizing optimal deformations, and PCA helps in dimension reduction and statistical modeling. One key challenge here is comparisons of shape graphs with vastly different complexities (in number of nodes and edges). This paper introduces a novel multi-scale representation of shape graphs to handle this challenge. Using the notions of (1) ``effective resistance" to cluster nodes and (2) elastic shape averaging of edge curves, one can reduce shape graph complexity while maintaining overall structures. This way, we can compare shape graphs by bringing them to similar complexity. We demonstrate these ideas on Retinal Blood Vessel (RBV) networks taken from the STARE and DRIVE databases.

stat.ME↗

Elastic Metrics on Spaces of Euclidean Curves: Theory and Algorithms

A main goal in the field of statistical shape analysis is to define computable and informative metrics on spaces of immersed manifolds, such as the space of curves in a Euclidean space. The approach taken in the elastic shape analysis framework is to define such a metric by starting with a reparameterization-invariant Riemannian metric on the space of parameterized shapes and inducing a metric on the quotient by the group of diffeomorphisms. This quotient metric is computed, in practice, by finding a registration of two shapes over the diffeomorphism group. For spaces of Euclidean curves, the initial Riemannian metric is frequently chosen from a two-parameter family of Sobolev metrics, called elastic metrics. Elastic metrics are especially convenient because, for several parameter choices, they are known to be locally isometric to Riemannian metrics for which one is able to solve the geodesic boundary problem explictly -- well-known examples of these local isometries include the complex square root transform of Younes, Michor, Mumford and Shah and square root velocity (SRV) transform of Srivastava, Klassen, Joshi and Jermyn. In this paper, we show that the SRV transform extends to elastic metrics for all choices of parameters, for curves in any dimension, thereby fully generalizing the work of many authors over the past two decades. We give a unified treatment of the elastic metrics: we extend results of Trouvé and Younes, Bruveris as well as Lahiri, Robinson and Klassen on the existence of solutions to the registration problem, we develop algorithms for computing distances and geodesics, and we apply these algorithms to metric learning problems, where we learn optimal elastic metric parameters for statistical shape analysis tasks.

math.DG↗

Fusion Frame Homotopy and Tightening Fusion Frames by Gradient Descent

Finite frames, or spanning sets for finite-dimensional Hilbert spaces, are a ubiquitous tool in signal processing. There has been much recent work on understanding the global structure of collections of finite frames with prescribed properties, such as spaces of unit norm tight frames. We extend some of these results to the more general setting of fusion frames -- a fusion frame is a collection of subspaces of a finite-dimensional Hilbert space with the property that any vector can be recovered from its list of projections. The notion of tightness extends to fusion frames, and we consider the following basic question: is the collection of tight fusion frames with prescribed subspace dimensions path connected? We answer (a generalization of) this question in the affirmative, extending the analogous result for unit norm tight frames proved by Cahill, Mixon and Strawn. We also extend a result of Benedetto and Fickus, who defined a natural functional on the space of unit norm frames (the frame potential), showed that its global minimizers are tight, and showed that it has no spurious local minimizers, meaning that gradient descent can be used to construct unit-norm tight frames. We prove the analogous result for the fusion frame potential of Casazza and Fickus, implying that, when tight fusion frames exist for a given choice of dimensions, they can be constructed via gradient descent. Our proofs use techniques from symplectic geometry and Mumford's geometric invariant theory.

math.FA↗

Distance distributions and inverse problems for metric measure spaces

Applications in data science, shape analysis and object classification frequently require comparison of probability distributions defined on different ambient spaces. To accomplish this, one requires a notion of distance on a given class of metric measure spaces -- that is, compact metric spaces endowed with probability measures. Such distances are typically defined as comparisons between metric measure space invariants, such as distance distributions (also referred to as shape distributions, distance histograms or shape contexts in the literature). Generally, distances defined in terms of distance distributions are actually pseudometrics, in that they may vanish when comparing nonisomorphic spaces. The goal of this paper is to set up a formal framework for assessing the discrimininative power of distance distributions, i.e., the extend to which these pseudometrics fail to define proper metrics. We formulate several precise inverse problems in terms of these invariants and answer them in several categories of metric measure spaces, including the category of plane curves, where we give a counterexample to the Curve Histogram Conjecture of Brinkman and Olver, the categories of embedded and Riemannian manifolds, where we obtain sphere rigidity results, and the category of metric graphs, where we obtain a local injectivity result along the lines of classical work of Boutin and Kemper on point cloud configurations. The inverse problems are further contextualized by the introduction of a variant of the Gromov-Wasserstein distance on the space of metric measure spaces, which is inspired by the original Monge formulation of optimal transport.

math.MG↗

Toric Symplectic Geometry and Full Spark Frames

The collection of $d \times N$ complex matrices with prescribed column norms and prescribed (nonzero) singular values forms a compact algebraic variety, which we refer to as a frame space. Elements of frame spaces -- i.e., frames -- are used to give robust representations of complex-valued signals, so that geometrical and measure-theoretic properties of frame spaces are of interest to the signal processing community. This paper is concerned with the following question: what is the probability that a frame drawn uniformly at random from a given frame space has the property that any subset of $d$ of its columns gives a basis for $\mathbb{C}^d$? We show that the probability is one, generalizing recent work of Cahill, Mixon and Strawn. To prove this, we first show that frame spaces are related to highly structured objects called toric symplectic manifolds. This relationship elucidates the geometric meaning of eigensteps -- certain spectral invariants of a frame -- and should be a more broadly applicable tool for studying probabilistic questions about the structure of frame spaces. As another application of our symplectic perspective, we completely characterize the norm and spectral data for which the corresponding frame space has singularities, answering some open questions in the frame theory literature.

math.FA↗