SearcharxivSearch

arXiv subjects

Samir Chowdhury

Publications and source records attributed to Samir Chowdhury.

17 recordsLinked to original sources

Hypergraph Co-Optimal Transport: Metric and Categorical Properties

Hypergraphs capture multi-way relationships in data, and they have consequently seen a number of applications in higher-order network analysis, computer vision, geometry processing, and machine learning. In this paper, we develop theoretical foundations for studying the space of hypergraphs using ingredients from optimal transport. By enriching a hypergraph with probability measures on its nodes and hyperedges, as well as relational information capturing local and global structures, we obtain a general and robust framework for studying the collection of all hypergraphs. First, we introduce a hypergraph distance based on the co-optimal transport framework of Redko et al. and study its theoretical properties. Second, we formalize common methods for transforming a hypergraph into a graph as maps between the space of hypergraphs and the space of graphs, and study their functorial properties and Lipschitz bounds. Finally, we demonstrate the versatility of our Hypergraph Co-Optimal Transport (HyperCOT) framework through various examples.

math.MG

Convergence of Hierarchical Clustering and Persistent Homology Methods on Directed Networks

While there has been much interest in adapting conventional clustering procedures---and in higher dimensions, persistent homology methods---to directed networks, little is known about the convergence of such methods. In order to even formulate the problem of convergence for such methods, one needs to stipulate a reasonable model for a directed network together with a flexible sampling theory for such a model. In this paper we propose and study a particular model of directed networks, and use this model to study the convergence of certain hierarchical clustering and persistent homology methods that accept any matrix of (possibly asymmetric) pairwise relations as input and produce dendrograms and persistence barcodes as outputs. We show that as points are sampled from some probability distribution, the output of each method converges almost surely to a dendrogram/barcode depending on the structure of the distribution.

cs.CG

The Metric Space of Networks

We study the question of reconstructing a weighted, directed network up to isomorphism from its motifs. In order to tackle this question we first relax the usual (strong) notion of graph isomorphism to obtain a relaxation that we call weak isomorphism. Then we identify a definition of distance on the space of all networks that is compatible with weak isomorphism. This global approach comes equipped with notions such as completeness, compactness, curves, and geodesics, which we explore throughout this paper. Furthermore, it admits global-to-local inference in the following sense: we prove that two networks are weakly isomorphic if and only if all their motif sets are identical, thus answering the network reconstruction question. Further exploiting the additional structure imposed by our network distance, we prove that two networks are weakly isomorphic if and only if certain essential associated structures---the skeleta of the respective networks---are strongly isomorphic.

cs.DM

Distances and Isomorphism between Networks: Stability and Convergence of Network Invariants

We develop the theoretical foundations of a generalized Gromov-Hausdorff distance between functions on networks that has recently been applied to various subfields of topological data analysis and optimal transport. These functional representations of networks, or networks for short, specialize in the finite setting to (possibly asymmetric) adjacency matrices and derived representations such as distance or kernel matrices. Existing literature utilizing these constructions cannot, however, benefit from continuous formulations because the continuum limits of finite networks under this distance are not well-understood. For example, while there are currently numerous persistent homology methods on finite networks, it is unclear if these methods produce well-defined persistence diagrams in the infinite setting. We resolve this situation by introducing the collection of compact networks that arises by taking continuum limits of finite networks and developing sampling results showing that this collection admits well-defined persistence diagrams. Compared to metric spaces, the isomorphism class of the generalized Gromov-Hausdorff distance over networks is rather complex, and contains representatives having different cardinalities and different topologies. We provide an exact characterization of a suitable notion of isomorphism for compact networks as well as alternative, stronger characterizations under additional topological regularity assumptions. Toward data applications, we describe a unified framework for developing quantitatively stable network invariants, provide basic examples, and cast existing results on the stability of persistent homology methods in this extended framework. To illustrate our theoretical results, we introduce a model of directed circles with finite reversibility and characterize their Dowker persistence diagrams.

cs.DM

Quantized Gromov-Wasserstein

The Gromov-Wasserstein (GW) framework adapts ideas from optimal transport to allow for the comparison of probability distributions defined on different metric spaces. Scalable computation of GW distances and associated matchings on graphs and point clouds have recently been made possible by state-of-the-art algorithms such as S-GWL and MREC. Each of these algorithmic breakthroughs relies on decomposing the underlying spaces into parts and performing matchings on these parts, adding recursion as needed. While very successful in practice, theoretical guarantees on such methods are limited. Inspired by recent advances in the theory of quantization for metric measure spaces, we define Quantized Gromov Wasserstein (qGW): a metric that treats parts as fundamental objects and fits into a hierarchy of theoretical upper bounds for the GW problem. This formulation motivates a new algorithm for approximating optimal GW matchings which yields algorithmic speedups and reductions in memory complexity. Consequently, we are able to go beyond outperforming state-of-the-art and apply GW matching at scales that are an order of magnitude larger than in the existing literature, including datasets containing over 1M points.

cs.LG

Generalized Spectral Clustering via Gromov-Wasserstein Learning

We establish a bridge between spectral clustering and Gromov-Wasserstein Learning (GWL), a recent optimal transport-based approach to graph partitioning. This connection both explains and improves upon the state-of-the-art performance of GWL. The Gromov-Wasserstein framework provides probabilistic correspondences between nodes of source and target graphs via a quadratic programming relaxation of the node matching problem. Our results utilize and connect the observations that the GW geometric structure remains valid for any rank-2 tensor, in particular the adjacency, distance, and various kernel matrices on graphs, and that the heat kernel outperforms the adjacency matrix in producing stable and informative node correspondences. Using the heat kernel in the GWL framework provides new multiscale graph comparisons without compromising theoretical guarantees, while immediately yielding improved empirical results. A key insight of the GWL framework toward graph partitioning was to compute GW correspondences from a source graph to a template graph with isolated, self-connected nodes. We show that when comparing against a two-node template graph using the heat kernel at the infinite time limit, the resulting partition agrees with the partition produced by the Fiedler vector. This in turn yields a new insight into the k-cut graph partitioning problem through the lens of optimal transport. Our experiments on a range of real-world networks achieve comparable results to, and in many cases outperform, the state-of-the-art achieved by GWL.

cs.LG

Path homology and temporal networks

We present an algorithm to compute path homology for simple digraphs, and use it to topologically analyze various small digraphs en route to an analysis of complex temporal networks which exhibit such digraphs as underlying motifs. The digraphs analyzed include all digraphs, directed acyclic graphs, and undirected graphs up to certain numbers of vertices, as well as some specially constructed cases. Using information from this analysis, we identify small digraphs contributing to path homology in dimension $2$ for three temporal networks, and relate these digraphs to network behavior. We conclude that path homology can provide insight into temporal network structure and vice versa.

cs.SI

Gromov-Wasserstein Averaging in a Riemannian Framework

We introduce a theoretical framework for performing statistical tasks---including, but not limited to, averaging and principal component analysis---on the space of (possibly asymmetric) matrices with arbitrary entries and sizes. This is carried out under the lens of the Gromov-Wasserstein (GW) distance, and our methods translate the Riemannian framework of GW distances developed by Sturm into practical, implementable tools for network data analysis. Our methods are illustrated on datasets of letter graphs, asymmetric stochastic blockmodel networks, and planar shapes viewed as metric spaces. On the theoretical front, we supplement the work of Sturm by producing additional results on the tangent structure of this "space of spaces", as well as on the gradient flow of the Fréchet functional on this space.

math.MG

Path homologies of deep feedforward networks

We provide a characterization of two types of directed homology for fully-connected, feedforward neural network architectures. These exact characterizations of the directed homology structure of a neural network architecture are the first of their kind. We show that the directed flag homology of deep networks reduces to computing the simplicial homology of the underlying undirected graph, which is explicitly given by Euler characteristic computations. We also show that the path homology of these networks is non-trivial in higher dimensions and depends on the number and size of the layers within the network. These results provide a foundation for investigating homological differences between neural network architectures and their realized structure as implied by their parameters.

math.AT

The Gromov-Wasserstein distance between networks and stable network invariants

We define a metric---the network Gromov-Wasserstein distance---on weighted, directed networks that is sensitive to the presence of outliers. In addition to proving its theoretical properties, we supply network invariants based on optimal transport that approximate this distance by means of lower bounds. We test these methods on a range of simulated network datasets and on a dataset of real-world global bilateral migration. For our simulations, we define a network generative model based on the stochastic block model. This may be of independent interest for benchmarking purposes.

cs.DM

Geodesics in persistence diagram space

It is known that for a variety of choices of metrics, including the standard bottleneck distance, the space of persistence diagrams admits geodesics. Typically these existence results produce geodesics that have the form of a convex combination. More specifically, given two persistence diagrams and a choice of metric, one obtains a bijection realizing the distance between the diagrams, and uses this bijection to linearly interpolate from one diagram to another. We prove that for several families of metrics, every geodesic in persistence diagram space arises as such a convex combination. For certain other choices of metrics, we explicitly construct infinite families of geodesics that cannot have this form.

math.MG

Vietoris-Rips Complexes of Regular Polygons

Persistent homology has emerged as a novel tool for data analysis in the past two decades. However, there are still very few shapes or even manifolds whose persistent homology barcodes (say of the Vietoris-Rips complex) are fully known. Towards this direction, let $P_n$ be the boundary of a regular polygon in the plane with $n$ sides; we describe the homotopy types of Vietoris-Rips complexes of $P_n$. Indeed, when $n=(k+1)!!$ is an odd double factorial, we provide a complete characterization of the homotopy types and persistent homology of the Vietoris-Rips complexes of $P_n$ up to a scale parameter $r_n$, where $r_n$ approaches the diameter of $P_n$ as $n\to\infty$. Surprisingly, these homotopy types include spheres of all dimensions. Roughly speaking, the number of higher-dimensional spheres appearing is linked to the number of equilateral (but not necessarily equiangular) stars that can be inscribed into $P_n$. As our main tool we use the recently-developed theory of cyclic graphs and winding fractions. Furthermore, we show that the Vietoris-Rips complex of an arbitrarily dense subset of $P_n$ need not be homotopy equivalent to the Vietoris-Rips complex of $P_n$ itself, and indeed, these two complexes can have different homology groups in arbitrarily high dimensions. As an application of our results, we provide a lower bound on the Gromov-Hausdorff distance between $P_n$ and the circle.

math.MG

Explicit Geodesics in Gromov-Hausdorff Space

We provide an alternative, constructive proof that the collection $\mathcal{M}$ of isometry classes of compact metric spaces endowed with the Gromov-Hausdorff distance is a geodesic space. The core of our proof is a construction of explicit geodesics on $\mathcal{M}$. We also provide several interesting examples of geodesics on $\mathcal{M}$, including a geodesic between $\mathbb{S}^0$ and $\mathbb{S}^n$ for any $n\geq 1$.

math.MG

A functorial Dowker theorem and persistent homology of asymmetric networks

We study two methods for computing network features with topological underpinnings: the Rips and Dowker persistent homology diagrams. Our formulations work for general networks, which may be asymmetric and may have any real number as an edge weight. We study the sensitivity of Dowker persistence diagrams to asymmetry via numerous theoretical examples, including a family of highly asymmetric cycle networks that have interesting connections to the existing literature. In particular, we characterize the Dowker persistence diagrams arising from asymmetric cycle networks. We investigate the stability properties of both the Dowker and Rips persistence diagrams, and use these observations to run a classification task on a dataset comprising simulated hippocampal networks. Our theoretical and experimental results suggest that Dowker persistence diagrams are particularly suitable for studying asymmetric networks. As a stepping stone for our constructions, we prove a functorial generalization of a theorem of Dowker, after whom our constructions are named.

math.AT

New families of simplicial filtration functors

The so called Čech and Vietoris-Rips simplicial filtrations are designed to capture information about the topological structure of metric datasets. These filtrations are two of the workhorses in the field of topological data analysis. They enjoy stability with respect to the Gromov-Hausdorff (GH) distance, and this stability property allows us to estimate the GH distance between finite metric space representations of the underlying datasets. Via the concept of Gromov's curvature sets we construct a rich theoretical framework of valuation-induced stable filtration functors. This framework includes the Čech and Vietoris-Rips filtration functors as well as many novel filtration functors that capture diverse features present in datasets. We further explore the concept of basepoint filtrations functors and use it to provide a classification of the filtration functors that we identify.

math.AT

The Importance of Forgetting: Limiting Memory Improves Recovery of Topological Characteristics from Neural Data

We develop of a line of work initiated by Curto and Itskov towards understanding the amount of information contained in the spike trains of hippocampal place cells via topology considerations. Previously, it was established that simply knowing which groups of place cells fire together in an animal's hippocampus is sufficient to extract the global topology of the animal's physical environment. We model a system where collections of place cells group and ungroup according to short-term plasticity rules. In particular, we obtain the surprising result that in experiments with spurious firing, the accuracy of the extracted topological information decreases with the persistence (beyond a certain regime) of the cell groups. This suggests that synaptic transience, or forgetting, is a mechanism by which the brain counteracts the effects of spurious place cell activity.

q-bio.NC

Persistent Path Homology of Directed Networks

While standard persistent homology has been successful in extracting information from metric datasets, its applicability to more general data, e.g. directed networks, is hindered by its natural insensitivity to asymmetry. We study a construction of homology of digraphs due to Grigor'yan, Lin, Muranov and Yau, and extend this construction to the persistent framework. The result, which we call persistent path homology, can provide information about the digraph structure of a directed network at varying resolutions. Moreover, this method encodes a rich level of detail about the asymmetric structure of the input directed network. We test our method on both simulated and real-world directed networks and conjecture some of its characteristics.

math.AT