Searcharxiv⌕ Search

arXiv subjects

Lori Ziegelmeier

Publications and source records attributed to Lori Ziegelmeier.

At least 19 recordsLinked to original sources

Topological summaries of fingerprint ridge patterns carry identity information

Fingerprints are the most widely deployed biometric. Verifying whether two impressions come from the same finger typically relies on minutiae, small landmarks such as skin ridge endings and bifurcations. These landmarks are extracted through a multi-stage pipeline of image enhancement, skeletonization, minutiae detection, and alignment. We investigate an alternative: using topological data analysis to represent the full pattern of skin ridges and valleys directly, bypassing minutiae detection and the downstream matching pipeline. We apply persistent homology, a topological tool that tracks how loops in the ridge pattern form and fill in across spatial scales, producing multi-scale summaries of ridge geometry. We develop and compare a range of verification methods on a standard benchmark dataset, FVC2000 DB1. Even the simplest topological summaries, with no trained parameters, substantially outperform geometry-only baselines. A trained method achieves an AUC of 0.91, while an optimal-transport method excels at the strictest false-accept thresholds, suggesting they capture different aspects of the ridge pattern. Fusing these two approaches yields the best performance at every low false-accept threshold we examine. Our results establish that these topological summaries capture substantial fingerprint identity information, far more effective for verification than raw pixel-level geometry. Because the entire pipeline is openly specified, it offers a transparent complement to minutiae-based systems, and we provide a modular framework for constructing, evaluating, and combining topological verification methods.

cs.CV↗

A U-match Algorithm for Persistent Relative Homology

A central problem in data-driven scientific inquiry is how to interpret structure in noisy, high-dimensional data. Topological data analysis (TDA) provides a solution via persistent homology, which encodes features of interest as topological holes within a filtration of data. The present work extends this framework to a related invariant which uncovers topological structure of a space relative to a subspace: persistent relative homology (PRH). We show that this invariant can be computed in a simple, highly transparent and general manner, using a two-step matrix reduction technique with worst-case time complexity comparable to ordinary persistent homology. We provide proofs demonstrating the correctness and computational complexity of this approach in addition to a performance-optimized implementation for a special case.

math.AT↗

A survey of simplicial, relative, and chain complex homology theories for hypergraphs

Hypergraphs have seen widespread application in network and data science communities in recent years. We present a survey of recent work to construct auxiliary structures from hypergraphs -- specifically simplicial, relative, and chain complexes -- that can be used to build homology theories for hypergraphs. We define and describe nine different constructions and their associated homology theories. We discuss some interesting properties of each homology theory to show how various hypergraph properties imply properties of the homology groups. We also include discussion of functoriality for several of the homology theories. Finally, we provide a series of illustrative examples by computing many of these homology theories for small hypergraphs to show the variability of the methods and build intuition.

math.AT↗

Higher-Order Network Structure Inference: A Topological Approach to Network Selection

Thresholding--the pruning of nodes or edges based on their properties or weights--is an essential preprocessing tool for extracting interpretable structure from complex network data, yet existing methods face several key limitations. Threshold selection often relies on heuristic methods or trial and error due to large parameter spaces and unclear optimization criteria, leading to sensitivity where small parameter variations produce significant changes in network structure. Moreover, most approaches focus on pairwise relationships between nodes, overlooking critical higher-order interactions involving three or more nodes. We introduce a systematic thresholding algorithm that leverages topological data analysis to identify optimal network parameters by accounting for higher-order structural relationships. Our method uses persistent homology to compute the stability of homological features across the parameter space, identifying parameter choices that are robust to small variations while preserving meaningful topological structure. Hyperparameters allow users to specify minimum requirements for topological features, effectively constraining the parameter search to avoid spurious solutions. We demonstrate the approach with an application in the Science of Science, where networks of scientific concepts are extracted from research paper abstracts, and concepts are connected when they co-appear in the same abstract. The flexibility of our approach allows researchers to incorporate domain-specific constraints and extends beyond network thresholding to general parameterization problems in data analysis.

cs.SI↗

Understanding U.S. Racial Segregation Through Persistent Homology

Racial segregation is a widespread social and physical phenomenon present in every city across the United States. Although prevalent nationwide, each city has a unique history of racial segregation, resulting in distinct "shapes" of segregation. We use persistent homology, a technique from applied algebraic topology, to investigate whether common patterns of racial segregation exist among U.S. cities. We explore two methods of constructing simplicial complexes that preserve geospatial data, applying them to White, Black, Asian, and Hispanic demographic data from the U.S. census for 112 U.S. cities. Using these methods, we cluster the cities based on their persistence to identify groups with similar segregation "shapes". Finally, we apply cluster analysis techniques to explore the characteristics of our clusters. This includes calculating the mean cluster statistics to gain insights into the demographics of each cluster and using the Adjusted Rand Index to compare our results with other clustering methods.

cs.SI↗

Image Triangulation Using the Sobel Operator for Vertex Selection

Image triangulation, the practice of decomposing images into triangles, deliberately employs simplification to create an abstracted representation. While triangulating an image is a relatively simple process, difficulties arise when determining which vertices produce recognizable and visually pleasing output images. With the goal of producing art, we discuss an image triangulation algorithm in Python that utilizes Sobel edge detection and point cloud sparsification to determine final vertices for a triangulation, resulting in the creation of artistic triangulated compositions.

cs.CG↗

A Topology Scavenger Hunt to Introduce Topological Data Analysis

Topology at the undergraduate level is often a theoretical mathematics course, introducing concepts from point-set topology or possibly algebraic topology. However, the last two decades have seen an explosion of growth in applied topology and topological data analysis, which are topics that can be presented in an accessible way to undergraduate students and can encourage exciting projects. For the past several years, the Topology course at Macalester College has included content from point-set and algebraic topology, as well as applied topology, culminating in a project chosen by the students. In the course, students work through a topology scavenger hunt as an activity to introduce the ideas and software behind some of the primary tools in topological data analysis, namely, persistent homology and mapper. This scavenger hunt includes a variety of point clouds of varying dimensions, such as an annulus in 2D, a bouquet of loops in 3D, a sphere in 4D, and a torus in 400D. The students' goal is to analyze each point cloud with a variety of software to infer the topological structure. After completing this activity, students are able to extend the ideas learned in the scavenger hunt to an open-ended capstone project. Examples of past projects include: using persistence to explore the relationship between country development and geography, to analyze congressional voting patterns, and to classify genres of a large corpus of texts by combining with tools from natural language processing and machine learning.

math.HO↗

Minimal Cycle Representatives in Persistent Homology using Linear Programming: an Empirical Study with User's Guide

Cycle representatives of persistent homology classes can be used to provide descriptions of topological features in data. However, the non-uniqueness of these representatives creates ambiguity and can lead to many different interpretations of the same set of classes. One approach to solving this problem is to optimize the choice of representative against some measure that is meaningful in the context of the data. In this work, we provide a study of the effectiveness and computational cost of several $\ell_1$-minimization optimization procedures for constructing homological cycle bases for persistent homology with rational coefficients in dimension one, including uniform-weighted and length-weighted edge-loss algorithms as well as uniform-weighted and area-weighted triangle-loss algorithms. We conduct these optimizations via standard linear programming methods, applying general-purpose solvers to optimize over column bases of simplicial boundary matrices. Our key findings are: (i) optimization is effective in reducing the size of cycle representatives, (ii) the computational cost of optimizing a basis of cycle representatives exceeds the cost of computing such a basis in most data sets we consider, (iii) the choice of linear solvers matters a lot to the computation time of optimizing cycles, (iv) the computation time of solving an integer program is not significantly longer than the computation time of solving a linear program for most of the cycle representatives, using the Gurobi linear solver, (v) strikingly, whether requiring integer solutions or not, we almost always obtain a solution with the same cost and almost all solutions found have entries in {-1, 0, 1} and therefore, are also solutions to a restricted $\ell_0$ optimization problem, and (vi) we obtain qualitatively different results for generators in Erdős-Rényi random clique complexes.

math.AT↗

U-match factorization: sparse homological algebra, lazy cycle representatives, and dualities in persistent (co)homology

Persistent homology is a leading tool in topological data analysis (TDA). Many problems in TDA can be solved via homological -- and indeed, linear -- algebra. However, matrices in this domain are typically large, with rows and columns numbered in billions. Low-rank approximation of such arrays typically destroys essential information; thus, new mathematical and computational paradigms are needed for very large, sparse matrices. We present the U-match matrix factorization scheme to address this challenge. U-match has two desirable features. First, it admits a compressed storage format that reduces the number of nonzero entries held in computer memory by one or more orders of magnitude over other common factorizations. Second, it permits direct solution of diverse problems in linear and homological algebra, without decompressing matrices stored in memory. These problems include look-up and retrieval of rows and columns; evaluation of birth/death times, and extraction of generators in persistent (co)homology; and, calculation of bases for boundary and cycle subspaces of filtered chain complexes. Such bases are key to unlocking a range of other topological techniques for use in TDA, and U-match factorization is designed to make such calculations broadly accessible to practitioners. As an application, we show that individual cycle representatives in persistent homology can be retrieved at time and memory costs orders of magnitude below current state of the art, via global duality. Moreover, the algebraic machinery needed to achieve this computation already exists in many modern solvers.

math.AT↗

Capturing Dynamics of Time-Varying Data via Topology

One approach to understanding complex data is to study its shape through the lens of algebraic topology. While the early development of topological data analysis focused primarily on static data, in recent years, theoretical and applied studies have turned to data that varies in time. A time-varying collection of metric spaces as formed, for example, by a moving school of fish or flock of birds, can contain a vast amount of information. There is often a need to simplify or summarize the dynamic behavior. We provide an introduction to topological summaries of time-varying metric spaces including vineyards [19], crocker plots [56], and multiparameter rank functions [37]. We then introduce a new tool to summarize time-varying metric spaces: a crocker stack. Crocker stacks are convenient for visualization, amenable to machine learning, and satisfy a desirable continuity property which we prove. We demonstrate the utility of crocker stacks for a parameter identification task involving an influential model of biological aggregations [58]. Altogether, we aim to bring the broader applied mathematics community up-to-date on topological summaries of time-varying metric spaces.

cs.LG↗

Analyzing Collective Motion with Machine Learning and Topology

We use topological data analysis and machine learning to study a seminal model of collective motion in biology [D'Orsogna et al., Phys. Rev. Lett. 96 (2006)]. This model describes agents interacting nonlinearly via attractive-repulsive social forces and gives rise to collective behaviors such as flocking and milling. To classify the emergent collective motion in a large library of numerical simulations and to recover model parameters from the simulation data, we apply machine learning techniques to two different types of input. First, we input time series of order parameters traditionally used in studies of collective motion. Second, we input measures based in topology that summarize the time-varying persistent homology of simulation data over multiple scales. This topological approach does not require prior knowledge of the expected patterns. For both unsupervised and supervised machine learning methods, the topological approach outperforms the one that is based on traditional order parameters.

math.AT↗

Local Versus Global Distances for Zigzag Persistence Modules

This short note establishes explicit and broadly applicable relationships between persistence-based distances computed locally and globally. In particular, we show that the bottleneck distance between two zigzag persistence modules restricted to an interval is always bounded above by the distance between the unrestricted versions. While this result is not surprising, it could have different practical implications. We give two related applications for metric graph distances, as well as an extension for the matching distance between multiparameter persistence modules.

math.AT↗

The Relationship Between the Intrinsic Cech and Persistence Distortion Distances for Metric Graphs

Metric graphs are meaningful objects for modeling complex structures that arise in many real-world applications, such as road networks, river systems, earthquake faults, blood vessels, and filamentary structures in galaxies. To study metric graphs in the context of comparison, we are interested in determining the relative discriminative capabilities of two topology-based distances between a pair of arbitrary finite metric graphs: the persistence distortion distance and the intrinsic Cech distance. We explicitly show how to compute the intrinsic Cech distance between two metric graphs based solely on knowledge of the shortest systems of loops for the graphs. Our main theorem establishes an inequality between the intrinsic Cech and persistence distortion distances in the case when one of the graphs is a bouquet graph and the other is arbitrary. The relationship also holds when both graphs are constructed via wedge sums of cycles and edges.

math.AT↗

Assessing biological models using topological data analysis

We use topological data analysis as a tool to analyze the fit of mathematical models to experimental data. This study is built on data obtained from motion tracking groups of aphids in [Nilsen et al., PLOS One, 2013] and two random walk models that were proposed to describe the data. One model incorporates social interactions between the insects, and the second model is a control model that excludes these interactions. We compare data from each model to data from experiment by performing statistical tests based on three different sets of measures. First, we use time series of order parameters commonly used in collective motion studies. These order parameters measure the overall polarization and angular momentum of the group, and do not rely on a priori knowledge of the models that produced the data. Second, we use order parameter time series that do rely on a priori knowledge, namely average distance to nearest neighbor and percentage of aphids moving. Third, we use computational persistent homology to calculate topological signatures of the data. Analysis of the a priori order parameters indicates that the interactive model better describes the experimental data than the control model does. The topological approach performs as well as these a priori order parameters and better than the other order parameters, suggesting the utility of the topological approach in the absence of specific knowledge of mechanisms underlying the data.

q-bio.QM↗

Metric reconstruction via optimal transport

Given a sample of points $X$ in a metric space $M$ and a scale $r>0$, the Vietoris-Rips simplicial complex $\mathrm{VR}(X;r)$ is a standard construction to attempt to recover $M$ from $X$ up to homotopy type. A deficiency of this approach is that $\mathrm{VR}(X;r)$ is not metrizable if it is not locally finite, and thus does not recover metric information about $M$. We attempt to remedy this shortcoming by defining a metric space thickening of $X$, which we call the \emph{Vietoris-Rips thickening} $\mathrm{VR}^m(X;r)$, via the theory of optimal transport. When $M$ is a complete Riemannian manifold, or alternatively a compact Hadamard space, we show that the the Vietoris-Rips thickening satisfies Hausmann's theorem ($\mathrm{VR}^m(M;r)\simeq M$ for $r$ sufficiently small) with a simpler proof: homotopy equivalence $\mathrm{VR}^m(M;r)\to M$ is canonically defined as a center of mass map, and its homotopy inverse is the (now continuous) inclusion map $M\hookrightarrow\mathrm{VR}^m(M;r)$. Furthermore, we describe the homotopy type of the Vietoris-Rips thickening of the $n$-sphere at the first positive scale parameter $r$ where the homotopy type changes.

math.MG↗

Mind the Gap: A Study in Global Development through Persistent Homology

The Gapminder project set out to use statistics to dispel simplistic notions about global development. In the same spirit, we use persistent homology, a technique from computational algebraic topology, to explore the relationship between country development and geography. For each country, four indicators, gross domestic product per capita; average life expectancy; infant mortality; and gross national income per capita, were used to quantify the development. Two analyses were performed. The first considers clusters of the countries based on these indicators, and the second uncovers cycles in the data when combined with geographic border structure. Our analysis is a multi-scale approach that reveals similarities and connections among countries at a variety of levels. We discover localized development patterns that are invisible in standard statistical methods.

math.AT↗

A Complete Characterization of the 1-Dimensional Intrinsic Cech Persistence Diagrams for Metric Graphs

Metric graphs are special types of metric spaces used to model and represent simple, ubiquitous, geometric relations in data such as biological networks, social networks, and road networks. We are interested in giving a qualitative description of metric graphs using topological summaries. In particular, we provide a complete characterization of the 1-dimensional intrinsic Cech persistence diagrams for metric graphs using persistent homology. Together with complementary results by Adamaszek et. al, which imply results on intrinsic Cech persistence diagrams in all dimensions for a single cycle, our results constitute important steps toward characterizing intrinsic Cech persistence diagrams for arbitrary metric graphs across all dimensions.

math.AT↗

Stratifying High Dimensional Data Based on Proximity to the Convex Hull Boundary

The convex hull of a set of points, $C$, serves to expose extremal properties of $C$ and can help identify elements in $C$ of high interest. For many problems, particularly in the presence of noise, the true vertex set (and facets) may be difficult to determine. One solution is to expand the list of high interest candidates to points lying near the boundary of the convex hull. We propose a quadratic program for the purpose of stratifying points in a data cloud based on proximity to the boundary of the convex hull. For each data point, a quadratic program is solved to determine an associated weight vector. We show that the weight vector encodes geometric information concerning the point's relationship to the boundary of the convex hull. The computation of the weight vectors can be carried out in parallel, and for a fixed number of points and fixed neighborhood size, the overall computational complexity of the algorithm grows linearly with dimension. As a consequence, meaningful computations can be completed on reasonably large, high dimensional data sets.

cs.CG↗