SearcharxivSearch

arXiv subjects

Thomas Weighill

Publications and source records attributed to Thomas Weighill.

At least 19 recordsLinked to original sources

Measuring Spatial Clustering via Metropolis-Hastings Diffusion Distance

We propose a novel measure of the discrepancy between two probability distributions $f$ and $g$ on a graph - which we call the diffusion distance - that measures the rate of convergence of $f$ to $g$ under a graph-constrained Markov chain with stationary distribution $g$. As a default choice for this Markov chain, we use the Metropolis-Hastings transition matrix targeting $g$ with proposals given by a random walk on the graph. Our primary case of interest is when the second distribution $g$ is uniform, in which case the diffusion distance becomes a measure of spatial clustering in $f$. Used in this way, (Metropolis-Hastings) diffusion distance to uniformity extends Moran's $I$-type measures of spatial autocorrelation by incorporating global graph geometry rather than just local patterns. Indeed, Moran's $I$, the most well-known measure of spatial autocorrelation, can be viewed as a one-step heuristic for diffusion distance, so long as specific spatial weights are used. We establish theoretical bounds and a stability result for our measure, connecting it to graph spectra and optimal transport. We then turn our attention to outlining a statistical test for spatial clustering using diffusion distance. Under permutation null models, we derive high-probability bounds on diffusion distance underpinned by exact spectral formulas for convergence of distributions, enabling an efficient statistical test for spatial clustering on large datasets. We empirically compare diffusion distance to Moran's $I$ both as a numerical measure and as a statistical test. We show that diffusion distance exhibits higher power on synthetic data using a stochastic block model. Empirical analysis of Black population distributions for 100 U.S. cities shows that diffusion distance detects subtle differences in urban segregation patterns that Moran's $I$ does not.

math.ST

Topological optimization with birth and death cochains

We introduce the notion of birth and death cochains as generalized versions of birth and death simplices in persistent cohomology. We show that birth and death cochains (unlike birth and death simplices) are always unique for a given persistent cohomology class. We use birth and death cochains to define birth and death content as generalizations of birth and death times. We then demonstrate the advantages of using that birth and death content as loss functions on a variety of topological optimization tasks with point clouds, time series and scalar fields. We close with a novel application of topological optimization to a dataset of arctic ice images.

math.AT

Through the Grapevine: Vineyard Distance as a Measure of Topological Dissimilarity

We introduce a new measure of distance between datasets, based on vineyards from topological data analysis, which we call the vineyard distance. Vineyard distance measures the extent of topological change along an interpolation from one dataset to another, either along a pre-computed trajectory or via a straight-line homotopy. We demonstrate through theoretical results and experiments that vineyard distance is less sensitive than $L^p$ distance (which considers every single data value), but more sensitive than Wasserstein distance between persistence diagrams (which accounts only for shape and not location). This allows vineyard distance to reveal distinctions that the other two distance measures cannot. In our paper, we establish theoretical results for vineyard distance including as upper and lower bounds. We then demonstrate the usefulness of vineyard distance on real-world data through applications to geospatial data and to neural network training dynamics.

math.AT

Classification of Firn Data via Topological Features

In this paper we evaluate the performance of topological features for generalizable and robust classification of firn image data, with the broader goal of understanding the advantages, pitfalls, and trade-offs in topological featurization. Firn refers to layers of granular snow within glaciers that haven't been compressed into ice. This compactification process imposes distinct topological and geometric structure on firn that varies with depth within the firn column, making topological data analysis (TDA) a natural choice for understanding the connection between depth and structure. We use two classes of topological features, sublevel set features and distance transform features, together with persistence curves, to predict sample depth from microCT images. A range of challenging training-test scenarios reveals that no one choice of method dominates in all categories, and uncoveres a web of trade-offs between accuracy, interpretability, and generalizability.

cs.CV

Lifting coarse homotopies

Coarse geometry, and in particular coarse homotopy theory, has proven to be a powerful tool for approaching problems in geometric group theory and higher index theory. In this paper, we continue to develop theory in this area by proving a Coarse Lifting Lemma with respect to a certain class of bornologous surjective maps. This class is wide enough to include quotients by coarsely discontinuous group actions, which allows us to obtain results concerning the coarse fundamental group of quotients which are analogous to classical topological results for the fundamental group. As an application, we compute the fundamental group of metric cones over negatively curved compact Riemannian manifolds.

math.GT

Generalized Dimension Reduction Using Semi-Relaxed Gromov-Wasserstein Distance

Dimension reduction techniques typically seek an embedding of a high-dimensional point cloud into a low-dimensional Euclidean space which optimally preserves the geometry of the input data. Based on expert knowledge, one may instead wish to embed the data into some other manifold or metric space in order to better reflect the geometry or topology of the point cloud. We propose a general method for manifold-valued multidimensional scaling based on concepts from optimal transport. In particular, we establish theoretical connections between the recently introduced semi-relaxed Gromov-Wasserstein (srGW) framework and multidimensional scaling by solving the Monge problem in this setting. We also derive novel connections between srGW distance and Gromov-Hausdorff distance. We apply our computational framework to analyze ensembles of political redistricting plans for states with two Congressional districts, achieving an effective visualization of the ensemble as a distribution on a circle which can be used to characterize typical neutral plans, and to flag outliers.

math.OC

Coarse embeddings of quotients by finite group actions

We prove that for a metric space $X$ and a finite group $G$ acting on $X$ by isometries, if $X$ coarsely embeds into a Hilbert space, then so does the quotient $X/G$. A crucial step towards our main result is to show that for any integer $k > 0$ the space of unordered $k$-tuples of points in Hilbert space, with the $1$-Wasserstein distance, itself coarsely embeds into Hilbert space. Our proof relies on establishing bounds on the sliced Wasserstein distance between empirical measures in $\mathbb{R}^n$.

math.MG

Coarse embeddability of Wasserstein space and the space of persistence diagrams

We prove an equivalence between open questions about the embeddability of the space of persistence diagrams and the space of probability distributions (i.e.~Wasserstein space). It is known that for many natural metrics, no coarse embedding of either of these two spaces into Hilbert space exists. Some cases remain open, however. In particular, whether coarse embeddings exist with respect to the $p$-Wasserstein distance for $1\leq p\leq 2$ remains an open question for the space of persistence diagrams and for Wasserstein space on the plane. In this paper, we show that embeddability for persistence diagrams \redd{is equivalent to} embeddability for Wasserstein space on $\mathbb{R}^2$. \redd{When $p > 1$, Wasserstein space on $\mathbb{R}^2$ is snowflake universal (an obstruction to embeddability into any Banach space of non-trivial type) if and only if the space of persistence diagrams is snowflake universal.

math.MG

Topological analysis of U.S. city demographics

We apply persistent homology, the main method in topological data analysis, to the study of demographic data. Persistence diagrams efficiently summarize information about clusters or peaks in a region's demographic data. To illustrate how persistence diagrams can be used for exploratory analysis, we undertake a study of the 100 largest U.S.~cities and their Black and Hispanic populations. We use our method to find clusters in individual cities, determine which cities are outliers and why, measure and describe change in demographic patterns over time, and roughly categorize cities into distinct groups based on the topology of their demographics. Along the way, we highlight the advantages and disadvantages of persistence diagrams as a tool for analyzing geospatial data.

math.AT

Reconfiguration of Polygonal Subdivisions via Recombination

Motivated by the problem of redistricting, we study area-preserving reconfigurations of connected subdivisions of a simple polygon. A connected subdivision of a polygon $\mathcal{R}$, called a district map, is a set of interior disjoint connected polygons called districts whose union equals $\mathcal{R}$. We consider the recombination as the reconfiguration move which takes a subdivision and produces another by merging two adjacent districts, and by splitting them into two connected polygons of the same area as the original districts. The complexity of a map is the number of vertices in the boundaries of its districts. Given two maps with $k$ districts, with complexity $O(n)$, and a perfect matching between districts of the same area in the two maps, we show constructively that $(\log n)^{O(\log k)}$ recombination moves are sufficient to reconfigure one into the other. We also show that $Ω(\log n)$ recombination moves are sometimes necessary even when $k=3$, thus providing a tight bound when $k=O(1)$.

cs.CG

Measuring Segregation via Analysis on Graphs

In this paper, we use analysis on graphs to study quantitative measures of segregation. We focus on a classical statistic from the geography and urban sociology literature known as Moran's I, which in our language is a score associated to a real-valued function on a graph, computed with respect to a spatial weight matrix such as the adjacency matrix associated to the geographic units that tile a city. Our results characterizing the extremal behavior of I illustrate the important role of the underlying graph structure, especially the degree distribution, in interpreting the score. In addition to the standard spatial weight matrices encoding unit adjacency, we consider the Laplacian L and a doubly-stochastic approximation M. These alternatives allow us to connect I to ideas from Fourier analysis and random walks. We offer illustrations of our theoretical results with a mix of stylized synthetic examples and real geographic/demographic data.

cs.SI

Geometric averages of partitioned datasets

We introduce a method for jointly registering ensembles of partitioned datasets in a way which is both geometrically coherent and partition-aware. Once such a registration has been defined, one can group partition blocks across datasets in order to extract summary statistics, generalizing the commonly used order statistics for scalar-valued data. By modeling a partitioned dataset as an unordered $k$-tuple of points in a Wasserstein space, we are able to draw from techniques in optimal transport. More generally, our method is developed using the formalism of local Fréchet means in symmetric products of metric spaces. We establish basic theory in this general setting, including Alexandrov curvature bounds and a verifiable characterization of local means. Our method is demonstrated on ensembles of political redistricting plans to extract and visualize basic properties of the space of plans for a particular state, using North Carolina as our main example.

cs.CG

Political Geography and Representation: A Case Study of Districting in Pennsylvania

This preprint offers a detailed look, both qualitative and quantitative, at districting with respect to recent voting patterns in one state: Pennsylvania. We investigate how much the partisan playing field is tilted by political geography. In particular we closely examine the role of scale. We find that partisan-neutral maps rarely give seats proportional to votes, and that making the district size smaller tends to make it even harder to find a proportional map. This preprint was prepared as a chapter in the forthcoming edited volume Political Geometry, an interdisciplinary collection of essays on redistricting. (mggg.org/gerrybook)

cs.CY

Three Applications of Entropy to Gerrymandering

This preprint is an exploration in how a single mathematical idea - entropy - can be applied to redistricting in a number of ways. It's meant to be read not so much as a call to action for entropy, but as a case study illustrating one of the many ways math can inform our thinking on redistricting problems. This preprint was prepared as a chapter in the forthcoming edited volume Political Geometry, an interdisciplinary collection of essays on redistricting. (mggg.org/gerrybook)

cs.IT

The (homological) persistence of gerrymandering

We apply persistent homology, the dominant tool from the field of topological data analysis, to study electoral redistricting. Our method combines the geographic information from a political districting plan with election data to produce a persistence diagram. We are then able to visualize and analyze large ensembles of computer-generated districting plans of the type commonly used in modern redistricting research (and court challenges). We set out three applications: zoning a state at each scale of districting, comparing elections, and seeking signals of gerrymandering. Our case studies focus on redistricting in Pennsylvania and North Carolina, two states whose legal challenges to enacted plans have raised considerable public interest in the last few years. To address the question of robustness of the persistence diagrams to perturbations in vote data and in district boundaries, we translate the classical stability theorem of Cohen--Steiner et al. into our setting and find that it can be phrased in a manner that is easy to interpret. We accompany the theoretical bound with an empirical demonstration to illustrate diagram stability in practice.

math.AT

Mal'tsev objects, $R_1$-spaces and ultrametric spaces

In this paper we introduce a notion of Mal'tsev object, and the dual notion of co-Mal'tsev object, in a general category. In particular, a category $\mathbb{C}$ is a Mal'tsev category if and only if every object in $\mathbb{C}$ is a Mal'tsev object. We show that for a well-powered regular category $\mathbb{C}$ which admits coproducts, the full subcategory of Mal'tsev objects is coreflective in $\mathbb{C}$. We show that the co-Mal'tsev objects in the category of topological spaces and continuous maps are precisely the $R_1$-spaces, and that the co-Mal'tsev objects in the category of metric spaces and short maps are precisely the ultrametric spaces.

math.CT

Geometry of Graph Partitions via Optimal Transport

We define a distance metric between partitions of a graph using machinery from optimal transport. Our metric is built from a linear assignment problem that matches partition components, with assignment cost proportional to transport distance over graph edges. We show that our distance can be computed using a single linear program without precomputing pairwise assignment costs and derive several theoretical properties of the metric. Finally, we provide experiments demonstrating these properties empirically, specifically focusing on its value for new problems in ensemble-based analysis of political districting plans.

math.OC

Extension theorems for large scale spaces via coarse neighbourhoods

We introduce the notion of (hybrid) large scale normal space and prove coarse geometric analogues of Urysohn's Lemma and the Tietze Extension Theorem for these spaces, where continuous maps are replaced by (continuous and) slowly oscillating maps. To do so, we first prove a general form of each of these results in the context of a set equipped with a neighbourhood operator satisfying certain axioms, from which we obtain both the classical topological results and the (hybrid) large scale results as corollaries. We prove that all metric spaces are hybrid large scale normal, and characterize those locally compact abelian groups which (as hybrid large scale spaces) are hybrid large scale normal. Finally, we look at some properties of the Higson compactifications and coronas of hybrid large scale normal spaces.

math.MG