SearcharxivSearch

arXiv subjects

Chad Giusti

Publications and source records attributed to Chad Giusti.

At least 19 recordsLinked to original sources

ALBATROSS: Cheap Filtration Based Geometry via Stochastic Sub-Sampling

Topological data analysis (TDA) detects geometric structure in biological data. However, many TDA algorithms are memory intensive and impractical for massive datasets. Here, we introduce a statistical protocol that reduces TDA's memory requirements and gives access to scientists with modest computing resources. We validate this protocol against two empirical datasets, showing that it replicates previous findings with much lower memory requirements. Finally, we demonstrate the power of the protocol by mapping the topology of functional correlations for the human cortex at high spatial resolution, something that was previously infeasible without this novel approach.

q-bio.QM

Dowker's theorem for higher-order relations

Given a relation $R \subseteq I \times J$ between two sets, Dowker's Theorem (1952) states that the homology groups of two associated simplicial complexes, now known as Dowker complexes, are isomorphic. In its modern form, the full result asserts a functorial homotopy equivalence between the two Dowker complexes. What can be said about relations defined on three or more sets? We present a simple generalization to multiway relations of the form $R \subseteq I_1 \times I_2 \times \cdots \times I_m$. The theorem asserts functorial homotopy equivalences between $m$ multiway Dowker complexes and a variant of the rectangle complex of Brun and Salbu from their recent short proof of Dowker's Theorem. Our proof uses Smale's homotopy mapping theorem and factors through a cellular Dowker lemma that expresses the main idea in more general form. To make the geometry more transparent, we work with a class of spaces called prod-complexes then transfer the results to simplicial complexes through a simplexification process. We conclude with a detailed study of ternary relations, identifying seven functorially defined homotopy types and twelve natural transformations between them.

math.AT

Tracking the topology of neural manifolds across populations

Neural manifolds summarize the intrinsic structure of the information encoded by a population of neurons. Advances in experimental techniques have made simultaneous recordings from multiple brain regions increasingly commonplace, raising the possibility of studying how these manifolds relate across populations. However, when the manifolds are nonlinear and possibly code for multiple unknown variables, it is challenging to extract robust and falsifiable information about their relationships. We introduce a framework, called the method of analogous cycles, for matching topological features of neural manifolds using only observed dissimilarity matrices within and between neural populations. We demonstrate via analysis of simulations and \emph{in vivo} experimental data that this method can be used to correctly identify multiple shared circular coordinate systems across both stimuli and inferred neural manifolds. Conversely, the method rejects matching features that are not intrinsic to one of the systems. Further, as this method is deterministic and does not rely on dimensionality reduction or optimization methods, it is amenable to direct mathematical investigation and interpretation in terms of the underlying neural activity. We thus propose the method of analogous cycles as a suitable foundation for a theory of cross-population analysis via neural manifolds.

q-bio.NC

Computing homology of $\mathbb{Z}_k$-complexes from their quotients

In this paper, we investigate the question of how one can recover the homology of a simplicial complex $X$ equipped with a regular action of a finite group $G$ from the structure of its quotient space $X/G.$ Specifically, we describe a process for enriching the structure of the chain complex $C_\ast(X/G; \mathbb{F})$ using the data of a complex of groups, a framework developed by Bridson and Corsen for encoding the local structure of a group action. We interpret this data through the lens of matrix representations of the acting group, and combine this structure with the standard simplicial boundary matrices for $X/G$ to construct a surrogate chain complex. In the case $G = \mathbb{Z}_k,$ the group ring $\mathbb{F}G$ is commutative and matrices over $\mathbb{F}G$ admit a Smith normal form, allowing us to recover the homology of $G$ from this surrogate complex. This algebraic approach complements the geometric compression algorithm for equivariant simplicial complexes described by Carbone, Nanda, and Naqvi.

math.AT

Eulerian magnitude homology: subgraph structure and random graphs

In this paper we explore the connection between the ranks of the magnitude homology groups of a graph and the structure of its subgraphs. To this end, we introduce variants of magnitude homology called eulerian magnitude homology and discriminant magnitude homology. Leveraging the combinatorics of the differential in magnitude homology, we illustrate a close relationship between the ranks of the eulerian magnitude homology groups on the first diagonal and counts of subgraphs which fall in specific classes. We leverage these tools to study limiting behavior of the eulerian magnitude homology groups for Erdos-Renyi random graphs and random geometric graphs, producing for both models a vanishing threshold for the eulerian magnitude homology groups on the first diagonal. This in turn provides a characterization of the generators for the corresponding magnitude homology groups. Finally, we develop an explicit asymptotic estimate the expected rank of eulerian magnitude homology along the first diagonal for these random graph models.

math.CO

A Topological Approach to Mapping Space Signatures

A common approach for describing classes of functions and probability measures on a topological space $\mathcal{X}$ is to construct a suitable map $Φ$ from $\mathcal{X}$ into a vector space, where linear methods can be applied to address both problems. The case where $\mathcal{X}$ is a space of paths $[0,1] \to \mathbb{R}^n$ and $Φ$ is the path signature map has received much attention in stochastic analysis and related fields. In this article we develop a generalized $Φ$ for the case where $\mathcal{X}$ is a space of maps $[0,1]^d \to \mathbb{R}^n$ for any $d \in \mathbb{N}$, and show that the map $Φ$ generalizes many of the desirable algebraic and analytic properties of the path signature to $d \ge 2$. The key ingredient to our approach is topological; in particular, our starting point is a generalisation of K-T Chen's path space cochain construction to the setting of cubical mapping spaces.

math.FA

Persistent Extension and Analogous Bars: Data-Induced Relations Between Persistence Barcodes

A central challenge in topological data analysis is the interpretation of barcodes. The classical algebraic-topological approach to interpreting homology classes is to build maps to spaces whose homology carries semantics we understand and then to appeal to functoriality. However, we often lack such maps in real data; instead, we must rely on a cross-dissimilarity measure between our observations of a system and a reference. In this paper, we develop a pair of computational homological algebra approaches for relating persistent homology classes and barcodes: persistent extension, which enumerates potential relations between cycles from two complexes built on the same vertex set, and the method of analogous bars, which utilizes persistent extension and the witness complex built from a cross-dissimilarity measure to provide relations across systems. We provide an implementation of these methods and demonstrate their use in comparing cycles between two samples from the same metric space and determining whether topology is maintained or destroyed under clustering and dimensionality reduction.

math.AT

Minimal Cycle Representatives in Persistent Homology using Linear Programming: an Empirical Study with User's Guide

Cycle representatives of persistent homology classes can be used to provide descriptions of topological features in data. However, the non-uniqueness of these representatives creates ambiguity and can lead to many different interpretations of the same set of classes. One approach to solving this problem is to optimize the choice of representative against some measure that is meaningful in the context of the data. In this work, we provide a study of the effectiveness and computational cost of several $\ell_1$-minimization optimization procedures for constructing homological cycle bases for persistent homology with rational coefficients in dimension one, including uniform-weighted and length-weighted edge-loss algorithms as well as uniform-weighted and area-weighted triangle-loss algorithms. We conduct these optimizations via standard linear programming methods, applying general-purpose solvers to optimize over column bases of simplicial boundary matrices. Our key findings are: (i) optimization is effective in reducing the size of cycle representatives, (ii) the computational cost of optimizing a basis of cycle representatives exceeds the cost of computing such a basis in most data sets we consider, (iii) the choice of linear solvers matters a lot to the computation time of optimizing cycles, (iv) the computation time of solving an integer program is not significantly longer than the computation time of solving a linear program for most of the cycle representatives, using the Gurobi linear solver, (v) strikingly, whether requiring integer solutions or not, we almost always obtain a solution with the same cost and almost all solutions found have entries in {-1, 0, 1} and therefore, are also solutions to a restricted $\ell_0$ optimization problem, and (vi) we obtain qualitatively different results for generators in Erdős-Rényi random clique complexes.

math.AT

U-match factorization: sparse homological algebra, lazy cycle representatives, and dualities in persistent (co)homology

Persistent homology is a leading tool in topological data analysis (TDA). Many problems in TDA can be solved via homological -- and indeed, linear -- algebra. However, matrices in this domain are typically large, with rows and columns numbered in billions. Low-rank approximation of such arrays typically destroys essential information; thus, new mathematical and computational paradigms are needed for very large, sparse matrices. We present the U-match matrix factorization scheme to address this challenge. U-match has two desirable features. First, it admits a compressed storage format that reduces the number of nonzero entries held in computer memory by one or more orders of magnitude over other common factorizations. Second, it permits direct solution of diverse problems in linear and homological algebra, without decompressing matrices stored in memory. These problems include look-up and retrieval of rows and columns; evaluation of birth/death times, and extraction of generators in persistent (co)homology; and, calculation of bases for boundary and cycle subspaces of filtered chain complexes. Such bases are key to unlocking a range of other topological techniques for use in TDA, and U-match factorization is designed to make such calculations broadly accessible to practitioners. As an application, we show that individual cycle representatives in persistent homology can be retrieved at time and memory costs orders of magnitude below current state of the art, via global duality. Moreover, the algebraic machinery needed to achieve this computation already exists in many modern solvers.

math.AT

Signatures, Lipschitz-free spaces, and paths of persistence diagrams

Paths of persistence diagrams provide a summary of the dynamic topological structure of a one-parameter family of metric spaces. These summaries can be used to study and characterize the dynamic shape of data such as swarming behavior in multi-agent systems, time-varying fMRI scans from neuroscience, and time-dependent scalar fields in hydrodynamics. While persistence diagrams can provide a powerful topological summary of data, the standard space of persistence diagrams lacks the sufficient algebraic and analytic structure required for many theoretical and computational analyses. We enrich the space of persistence diagrams by isometrically embedding it into a Lipschitz-free space, a Banach space built from a universal construction. We utilize the Banach space structure to define bounded variation paths of persistence diagrams, which can be studied using the path signature, a reparametrization-invariant characterization of paths valued in a Banach space. The signature is universal and characteristic which allows us to theoretically characterize measures on the space of paths and motivates its use in the context of kernel methods. However, kernel methods often require a feature map into a Hilbert space, so we introduce the moment map, a stable and injective feature map for static persistence diagrams, and compose it with the discrete path signature, producing a computable feature map into a Hilbert space. Finally, we demonstrate the efficacy of our methods by applying this to a parameter estimation problem for a 3D model of swarming behavior.

math.AT

Mod-two cohomology rings of alternating groups

We calculate the mod-two cohomology of all alternating groups together, with both cup and transfer product structures, which in particular determines the additive structure and ring structure of the cohomology of individual groups. We show that there are no nilpotent elements in the cohomology rings of individual alternating groups. We calculate the action of the Steenrod algebra and discuss individual component rings. A range of techniques is needed: an almost Hopf ring structure associated to the embeddings of products of alternating groups, the Gysin sequence relating the cohomology of alternating groups to that of symmetric groups, Fox-Neuwirth resolutions, and restriction to elementary abelian subgroups.

math.AT

Iterated Integrals and Population Time Series Analysis

One of the core advantages topological methods for data analysis provide is that the language of (co)chains can be mapped onto the semantics of the data, providing a natural avenue for human understanding of the results. Here, we describe such a semantic structure on Chen's classical iterated integral cochain model for paths in Euclidean space. Specifically, in the context of population time series data, we observe that iterated integrals provide a model-free measure of pairwise influence that can be used for causality inference. Along the way, we survey recent results and applications, review the current standard methods for causality inference, and briefly provide our outlook on generalizations to go beyond time series data.

stat.OT

Effective learning is accompanied by high dimensional and efficient representations of neural activity

A fundamental cognitive process is the ability to map value and identity onto objects as we learn about them. Exactly how such mental constructs emerge and what kind of space best embeds this mapping remains incompletely understood. Here we develop tools to quantify the space and organization of such a mapping, thereby providing a framework for studying the geometric representations of neural responses as reflected in functional MRI. Considering how human subjects learn the values of novel objects, we show that quick learners have a higher dimensional geometric representation than slow learners, and hence more easily distinguishable whole-brain responses to objects of different value. Furthermore, we find that quick learners display a more compact embedding of their neural responses and hence have a higher ratio of their task-based dimension to their embedding dimension -- consistent with a greater efficiency of cognitive coding. Lastly, we investigate the neurophysiological drivers of high dimensional patterns at both regional and voxel levels, and we complete our study with a complementary test of the distinguishability of associated whole-brain responses. Our results demonstrate a spatial organization of neural responses characteristic of learning, and offer a suite of geometric measures applicable to the study of efficient coding in higher-order cognitive processes more broadly.

q-bio.NC

Developmental increases in white matter network controllability support a growing diversity of brain dynamics

As the human brain develops, it increasingly supports coordinated control of neural activity. The mechanism by which white matter evolves to support this coordination is not well understood. We use a network representation of diffusion imaging data from 882 youth ages 8 to 22 to show that white matter connectivity becomes increasingly optimized for a diverse range of predicted dynamics in development. Notably, stable controllers in subcortical areas are negatively related to cognitive performance. Investigating structural mechanisms supporting these changes, we simulate network evolution with a set of growth rules. We find that all brain networks are structured in a manner highly optimized for network control, with distinct control mechanisms predicted in child versus older youth. We demonstrate that our results cannot be simply explained by changes in network modularity. This work reveals a possible mechanism of human brain development that preferentially optimizes dynamic network control over static network architecture.

q-bio.NC

Knowledge gaps in the early growth of semantic networks

Understanding the features of and mechanisms behind language learning can provide insights into the general process of knowledge acquisition. Recent methods from network science applied to language learning have advanced the field, particularly by noting associations between densely connected words and acquisition. However, the importance of sparse areas of the network, or knowledge gaps, remains unexplored. Here we create a semantic feature network in which words correspond to nodes and in which connections correspond to semantic similarity. We develop a new analytical approach built on principles of applied topology to query the prevalence of knowledge gaps, which we propose manifest as cavities within the network. We detect topological cavities of multiple dimensions in the growing semantic feature network of children ages 16 to 30 months. The pattern of cavity appearance matches that of a constrained null model, created by predefining the affinity of each node for connections. Furthermore, when word acquisition time is computed from children of mothers with differing levels of education, we find that despite variation at the word level, the global organization as measured by persistent homology remains comparable. We show that topological properties of a node correlate with filling in cavities better than simple lexical properties such as the length and frequency of the corresponding word. Finally, we show that the large-scale architecture of the semantic feature network is topologically accommodating to many node orders. We discuss the importance of topology in language learning, and we speculate that the formation and filling of knowledge gaps may be a robust feature of knowledge acquisition.

q-bio.NC

On open and closed convex codes

Neural codes serve as a language for neurons in the brain. Convex codes, which arise from the pattern of intersections of convex sets in Euclidean space, are of particular relevance to neuroscience. Not every code is convex, however, and the combinatorial properties of a code that determine its convexity are still poorly understood. Here we find that a code that can be realized by a collection of open convex sets may or may not be realizable by closed convex sets, and vice versa, establishing that open convex and closed convex codes are distinct classes. We also prove that max intersection-complete codes (i.e. codes that contain all intersections of maximal codewords) are both open convex and closed convex, and provide an upper bound for their minimal embedding dimension. Finally, we show that the addition of non-maximal codewords to an open convex code preserves convexity.

math.CO

Cliques and Cavities in the Human Connectome

Encoding brain regions and their connections as a network of nodes and edges captures many of the possible paths along which information can be transmitted as humans process and perform complex behaviors. Because cognitive processes involve large and distributed networks of brain areas, examinations of multi-node routes within larger connection patterns can offer fundamental insights into the complexities of brain function. Here, we investigate both densely connected groups of nodes that could perform local computations as well as larger patterns of interactions that would allow for parallel processing. Finding such structures necessitates we move from considering pairwise interactions to capturing higher order relations, concepts naturally expressed in the language of algebraic topology. These tools can be used to study mesoscale structures arising from the arrangement of densely connected substructures called cliques in otherwise sparsely connected brain networks. We detect cliques (all-to-all connected sets of brain regions) in the average structural connectomes of 8 healthy adults and discover the presence of more large cliques than expected in null networks constructed via wiring minimization, providing architecture through which brain network can perform rapid, local processing. We then locate topological cavities of different dimensions, around which information may flow in either diverging or converging patterns. These cavities exist consistently across subjects, differ from those observed in null model networks, and link regions of early and late evolutionary origin in long loops, underscoring their unique role in controlling brain function. These results offer a first demonstration that techniques from algebraic topology offer a novel perspective on structural connectomics, highlighting loop-like paths as crucial features in the human brain's structural architecture.

q-bio.NC

Topological and geometric measurements of force chain structure

Developing quantitative methods for characterizing structural properties of force chains in densely packed granular media is an important step toward understanding or predicting large-scale physical properties of a packing. A promising framework in which to develop such methods is network science, which can be used to translate particle locations and force contacts to a graph in which particles are represented by nodes and forces between particles are represented by weighted edges. Applying network-based community-detection techniques to extract force chains opens the door to developing statistics of force chain structure, with the goal of identifying shape differences across packings, and providing a foundation on which to build predictions of bulk material properties from mesoscale network features. Here, we discuss a trio of related but fundamentally distinct measurements of mesoscale structure of force chains in arbitrary 2D packings, including a novel statistic derived using tools from algebraic topology, which together provide a tool set for the analysis of force chain architecture. We demonstrate the utility of this tool set by detecting variations in force chain architecture with pressure. Collectively, these techniques can be generalized to 3D packings, and to the assessment of continuous deformations of packings under stress or strain.

cond-mat.soft