SearcharxivSearch

arXiv subjects

Donald R. Sheehy

Publications and source records attributed to Donald R. Sheehy.

17 recordsLinked to original sources

Product Range Search Problem

Given a metric space, a standard metric range search, given a query (q, r), finds all points within distance r of the point q. Suppose now we have two different metrics d1 and d2. A product range query (q, r1, r2) is a point q and two radii $r1$ and $r2$. The output is all points within distance $r1$ of q with respect to d1 and all points within $r2$ of q with respect to $d2$. In other words, it is the intersection of two searches. We present two data structures for approximate product range search in doubling metrics. Both data structures use a net-tree variant, the greedy tree. The greedy tree is a data structure that can efficiently answer approximate range searches in doubling metrics. The first data structure is a generalization of the range tree from computational geometry using greedy trees rather than binary trees. The second data structure is a single greedy tree constructed on the product induced by the two metrics.

cs.CG

Approximating the Directed Hausdorff Distance

The Hausdorff distance is a metric commonly used to compute the set similarity of geometric sets. For sets containing a total of $n$ points, the exact distance can be computed naïvely in $O(n^2)$ time. In this paper, we show how to preprocess point sets individually so that the Hausdorff distance of any pair can then be approximated in linear time. We assume that the metric is doubling. The preprocessing time for each set is $O(n\log Δ)$ where $Δ$ is the ratio of the largest to smallest pairwise distances of the input. In theory, this can be reduced to $O(n\log n)$ time using a much more complicated algorithm. We compute $(1+\varepsilon)$-approximate Hausdorff distance in $(2 + \frac{1}{\varepsilon})^{O(d)}n$ time in a metric space with doubling dimension $d$. The $k$-partial Hausdorff distance ignores $k$ outliers to increase stability. Additionally, we give a linear-time algorithm to compute directed $k$-partial Hausdorff distance for all values of $k$ at once with no change to the preprocessing.

cs.CG

A Theory of Sub-Barcodes

From the work of Bauer and Lesnick, it is known that there is no functor from the category of pointwise finite-dimensional persistence modules to the category of barcodes and overlap matchings. In this work, we introduce sub-barcodes and show that there is a functor from the category of factorizations of persistence module homomorphisms to a poset of barcodes ordered by the sub-barcode relation. Sub-barcodes and factorizations provide a looser alternative to bottleneck matchings and interleavings that can give strong guarantees in a number of settings that arise naturally in topological data analysis. The main use of sub-barcodes is to make strong claims about an unknown barcode in the absence of an interleaving. For example, given only upper and lower bounds $g\geq f\geq \ell$ of an unknown real-valued function $f$, a sub-barcode associated with $f$ can be constructed from $\ell$ and $g$ alone. We propose a theory of sub-barcodes and observe that the subobjects in the category of functors from intervals to matchings naturally correspond to sub-barcodes.

cs.CG

A Sparse Delaunay Filtration

We show how a filtration of Delaunay complexes can be used to approximate the persistence diagram of the distance to a point set in $R^d$. Whereas the full Delaunay complex can be used to compute this persistence diagram exactly, it may have size $O(n^{\lceil d/2 \rceil})$. In contrast, our construction uses only $O(n)$ simplices. The central idea is to connect Delaunay complexes on progressively denser subsamples by considering the flips in an incremental construction as simplices in $d+1$ dimensions. This approach leads to a very simple and straightforward proof of correctness in geometric terms, because the final filtration is dual to a $(d+1)$-dimensional Voronoi construction similar to the standard Delaunay filtration complex. We also, show how this complex can be efficiently constructed.

cs.CG

Sketching Persistence Diagrams

Given a persistence diagram with $n$ points, we give an algorithm that produces a sequence of $n$ persistence diagrams converging in bottleneck distance to the input diagram, the $i$th of which has $i$ distinct (weighted) points and is a $2$-approximation to the closest persistence diagram with that many distinct points. For each approximation, we precompute the optimal matching between the $i$th and the $(i+1)$st. Perhaps surprisingly, the entire sequence of diagrams as well as the sequence of matchings can be represented in $O(n)$ space. The main approach is to use a variation of the greedy permutation of the persistence diagram to give good Hausdorff approximations and assign weights to these subsets. We give a new algorithm to efficiently compute this permutation, despite the high implicit dimension of points in a persistence diagram due to the effect of the diagonal. The sketches are also structured to permit fast (linear time) approximations to the Hausdorff distance between diagrams -- a lower bound on the bottleneck distance. For approximating the bottleneck distance, sketches can also be used to compute a linear-size neighborhood graph directly, obviating the need for geometric data structures used in state-of-the-art methods for bottleneck computation.

cs.CG

An Efficient Algorithm for Topological Characterisation of Worm-Like and Branched Micelle Structures from Simulations

Many surfactant-based formulations are utilised in industry as they produce desirable visco-elastic properties at low-concentrations. These properties are due to the presence of worm-like micelles (WLM) and, as a result, understanding the processes that lead to WLM formation is of significant interest. Various experimental techniques have been applied with some success to this problem but can encounter issues probing key microscopic characteristics or the specific regimes of interest. The complementary use of computer simulations could provide an alternate route to accessing their structural and dynamic behaviour. However, few computational methods exist for measuring key characteristics of WLMs formed in particle simulations. Further, their mathematical formulation are challenged by WLMs with sharp curvature profiles or density fluctuations along the backbone. Here we present a new topological algorithm for identifying and characterising WLMs micelles in particle simulations which has desirable mathematical properties that address short-comings in previous techniques. We apply the algorithm to the case of Sodium dodecyl sulfate (SDS) micelles to demonstrate how it can be used to construct a comprehensive topological characterisation of the observed structures.

cond-mat.soft

Randomized Incremental Construction of Net-Trees

Net-trees are a general purpose data structure for metric data that have been used to solve a wide range of algorithmic problems. We give a simple randomized algorithm to construct net-trees on doubling metrics using $O(n\log n)$ time in expectation. Along the way, we define a new, linear-size net-tree variant that simplifies the analyses and algorithms. We show a connection between these trees and approximate Voronoi diagrams and use this to simplify the point location necessary in net-tree construction. Our analysis uses a novel backwards analysis that may be of independent interest.

cs.CG

Adaptive Metrics for Adaptive Samples

In this paper we consider adaptive sampling's local-feature size, used in surface reconstruction and geometric inference, with respect to an arbitrary landmark set rather than the medial axis and relate it to a path-based adaptive metric on Euclidean space. We prove a near-duality between adaptive samples in the Euclidean metric space and uniform samples in this alternate metric space which results in topological interleavings between the offsets generated by this metric and those generated by an linear approximation of it. After smoothing the distance function associated to the adaptive metric, we apply a result from the theory of critical points of distance functions to the interleaved spaces which yields a computable homology inference scheme assuming one has Hausdorff-close samples of the domain and the landmark set.

cs.CG

The Generalized Persistent Nerve Theorem

In this paper a parameterized generalization of a good cover filtration is introduced called an ε-good cover, defined as a cover filtration in which the reduced homology groups of the image of the inclusions between the intersections of the cover filtration at two scales ε apart are trivial. Assuming that one has an ε-good cover filtration of a finite simplicial filtration, we prove a tight bound on the bottleneck distance between the persistence diagrams of the nerve filtration and the simplicial filtration that is linear with respect to ε and the homology dimension. This bound is the result of a computable chain map from the nerve filtration to the space filtration's chain complexes at a further scale. Quantitative guarantees for covers that are not good are useful for when one is working a non-convex metric space, or one has more simplicial covers that are not the result of triangulations of metric balls. The Persistent Nerve Lemma is also a corollary of our theorem as good covers are 0-good covers. Lastly, a technique is introduced that symmetrizes the asymmetric interleaving used to prove the bound by shifting the nerve filtration's persistence module, improving the interleaving constant by a factor of 2.

math.AT

Supporting Ruled Polygons

We explore several problems related to ruled polygons. Given a ruling of a polygon $P$, we consider the Reeb graph of $P$ induced by the ruling. We define the Reeb complexity of $P$, which roughly equates to the minimum number of points necessary to support $P$. We give asymptotically tight bounds on the Reeb complexity that are also tight up to a small additive constant. When restricted to the set of parallel rulings, we show that the Reeb complexity can be computed in polynomial time.

cs.CG

Approximating the Simplicial Depth

Let $P$ be a set of $n$ points in $d$-dimensions. The simplicial depth, $σ_P(q)$ of a point $q$ is the number of $d$-simplices with vertices in $P$ that contain $q$ in their convex hulls. The simplicial depth is a notion of data depth with many applications in robust statistics and computational geometry. Computing the simplicial depth of a point is known to be a challenging problem. The trivial solution requires $O(n^{d+1})$ time whereas it is generally believed that one cannot do better than $O(n^{d-1})$. In this paper, we consider approximation algorithms for computing the simplicial depth of a point. For $d=2$, we present a new data structure that can approximate the simplicial depth in polylogarithmic time, using polylogarithmic query time. In 3D, we can approximate the simplicial depth of a given point in near-linear time, which is clearly optimal up to polylogarithmic factors. For higher dimensions, we consider two approximation algorithms with different worst-case scenarios. By combining these approaches, we compute a $(1+\varepsilon)$-approximation of the simplicial depth in time $\tilde{O}(n^{d/2 + 1})$ ignoring polylogarithmic factor. All of these algorithms are Monte Carlo algorithms. Furthermore, we present a simple strategy to compute the simplicial depth exactly in $O(n^d \log n)$ time, which provides the first improvement over the trivial $O(n^{d+1})$ time algorithm for $d>4$. Finally, we show that computing the simplicial depth exactly is #P-complete and W[1]-hard if the dimension is part of the input.

cs.CG

A Geometric Perspective on Sparse Filtrations

We present a geometric perspective on sparse filtrations used in topological data analysis. This new perspective leads to much simpler proofs, while also being more general, applying equally to Rips filtrations and Cech filtrations for any convex metric. We also give an algorithm for finding the simplices in such a filtration and prove that the vertex removal can be implemented as a sequence of elementary edge collapses.

cs.CG

Approximating Nearest Neighbor Distances

Several researchers proposed using non-Euclidean metrics on point sets in Euclidean space for clustering noisy data. Almost always, a distance function is desired that recognizes the closeness of the points in the same cluster, even if the Euclidean cluster diameter is large. Therefore, it is preferred to assign smaller costs to the paths that stay close to the input points. In this paper, we consider the most natural metric with this property, which we call the nearest neighbor metric. Given a point set P and a path $γ$, our metric charges each point of $γ$ with its distance to P. The total charge along $γ$ determines its nearest neighbor length, which is formally defined as the integral of the distance to the input points along the curve. We describe a $(3+\varepsilon)$-approximation algorithm and a $(1+\varepsilon)$-approximation algorithm to compute the nearest neighbor metric. Both approximation algorithms work in near-linear time. The former uses shortest paths on a sparse graph using only the input points. The latter uses a sparse sample of the ambient space, to find good approximate geodesic paths.

cs.CG

Efficient and Robust Persistent Homology for Measures

We extend the notion of the distance to a measure from Euclidean space to probability measures on general metric spaces as a way to do topological data analysis in a way that is robust to noise and outliers. We then give an efficient way to approximate the sub-level sets of this function by a union of metric balls and extend previous results on sparse Rips filtrations to this setting. This robust and efficient approach to topological data analysis is illustrated with several examples from an implementation.

cs.CG

A New Approach to Output-Sensitive Voronoi Diagrams and Delaunay Triangulations

We describe a new algorithm for computing the Voronoi diagram of a set of $n$ points in constant-dimensional Euclidean space. The running time of our algorithm is $O(f \log n \log Δ)$ where $f$ is the output complexity of the Voronoi diagram and $Δ$ is the spread of the input, the ratio of largest to smallest pairwise distances. Despite the simplicity of the algorithm and its analysis, it improves on the state of the art for all inputs with polynomial spread and near-linear output size. The key idea is to first build the Voronoi diagram of a superset of the input points using ideas from Voronoi refinement mesh generation. Then, the extra points are removed in a straightforward way that allows the total work to be bounded in terms of the output complexity, yielding the output sensitive bound. The removal only involves local flips and is inspired by kinetic data structures.

cs.CG

A Fast Algorithm for Well-Spaced Points and Approximate Delaunay Graphs

We present a new algorithm that produces a well-spaced superset of points conforming to a given input set in any dimension with guaranteed optimal output size. We also provide an approximate Delaunay graph on the output points. Our algorithm runs in expected time $O(2^{O(d)}(n\log n + m))$, where $n$ is the input size, $m$ is the output point set size, and $d$ is the ambient dimension. The constants only depend on the desired element quality bounds. To gain this new efficiency, the algorithm approximately maintains the Voronoi diagram of the current set of points by storing a superset of the Delaunay neighbors of each point. By retaining quality of the Voronoi diagram and avoiding the storage of the full Voronoi diagram, a simple exponential dependence on $d$ is obtained in the running time. Thus, if one only wants the approximate neighbors structure of a refined Delaunay mesh conforming to a set of input points, the algorithm will return a size $2^{O(d)}m$ graph in $2^{O(d)}(n\log n + m)$ expected time. If $m$ is superlinear in $n$, then we can produce a hierarchically well-spaced superset of size $2^{O(d)}n$ in $2^{O(d)}n\log n$ expected time.

cs.CG

Linear-Size Approximations to the Vietoris-Rips Filtration

The Vietoris-Rips filtration is a versatile tool in topological data analysis. It is a sequence of simplicial complexes built on a metric space to add topological structure to an otherwise disconnected set of points. It is widely used because it encodes useful information about the topology of the underlying metric space. This information is often extracted from its so-called persistence diagram. Unfortunately, this filtration is often too large to construct in full. We show how to construct an O(n)-size filtered simplicial complex on an $n$-point metric space such that its persistence diagram is a good approximation to that of the Vietoris-Rips filtration. This new filtration can be constructed in $O(n\log n)$ time. The constant factors in both the size and the running time depend only on the doubling dimension of the metric space and the desired tightness of the approximation. For the first time, this makes it computationally tractable to approximate the persistence diagram of the Vietoris-Rips filtration across all scales for large data sets. We describe two different sparse filtrations. The first is a zigzag filtration that removes points as the scale increases. The second is a (non-zigzag) filtration that yields the same persistence diagram. Both methods are based on a hierarchical net-tree and yield the same guarantees.

cs.CG