SearcharxivSearch

arXiv subjects

Omer Gold

Publications and source records attributed to Omer Gold.

5 recordsLinked to original sources

Dynamic Time Warping and Geometric Edit Distance: Breaking the Quadratic Barrier

Dynamic Time Warping (DTW) and Geometric Edit Distance (GED) are basic similarity measures between curves or general temporal sequences (e.g., time series) that are represented as sequences of points in some metric space $(X, \mathrm{dist})$. The DTW and GED measures are massively used in various fields of computer science, computational biology, and engineering. Consequently, the tasks of computing these measures are among the core problems in P. Despite extensive efforts to find more efficient algorithms, the best-known algorithms for computing the DTW or GED between two sequences of points in $X = \mathbb{R}^d$ are long-standing dynamic programming algorithms that require quadratic runtime, even for the one-dimensional case $d = 1$, which is perhaps one of the most used in practice. In this paper, we break the nearly 50 years old quadratic time bound for computing DTW or GED between two sequences of $n$ points in $\mathbb{R}$, by presenting deterministic algorithms that run in $O\left( n^2 / \log\log n \right)$ time. Our algorithms can be extended to work also for higher dimensional spaces $\mathbb{R}^d$, for any constant $d$, when the underlying distance-metric $\mathrm{dist}$ is polyhedral (e.g., $L_1, L_\infty$).

cs.DS

Diameter Spanners, Eccentricity Spanners, and Approximating Extremal Distances

The diameter of a graph is one if its most important parameters, being used in many real-word applications. In particular, the diameter dictates how fast information can spread throughout data and communication networks. Thus, it is a natural question to ask how much can we sparsify a graph and still guarantee that its diameter remains preserved within an approximation $t$. This property is captured by the notion of extremal-distance spanners. Given a graph $G=(V,E)$, a subgraph $H=(V,E_H)$ is defined to be a $t$-diameter spanner if the diameter of $H$ is at most $t$ times the diameter of $G$. We show that for any $n$-vertex and $m$-edges directed graph $G$, we can compute a sparse subgraph $H$ that is a $(1.5)$-diameter spanner of $G$, such that $H$ contains at most $\tilde O(n^{1.5})$ edges. We also show that the stretch factor cannot be improved to $(1.5-ε)$. For a graph whose diameter is bounded by some constant, we show the existence of $\frac{5}{3}$-diameter spanner that contains at most $\tilde O(n^\frac{4}{3})$ edges. We also show that this bound is tight. Additionally, we present other types of extremal-distance spanners, such as $2$-eccentricity spanners and $2$-radius spanners, both contain only $\tilde O(n)$ edges and are computable in $\tilde O(m)$ time. Finally, we study extremal-distance spanners in the dynamic and fault-tolerant settings. An interesting implication of our work is the first $\tilde O(m)$-time algorithm for computing $2$-approximation of vertex eccentricities in general directed weighted graphs. Backurs et al. [STOC 2018] gave an $\tilde O(m\sqrt{n})$ time algorithm for this problem, and also showed that no $O(n^{2-o(1)})$ time algorithm can achieve an approximation factor better than $2$ for graph eccentricities, unless SETH fails; this shows that our approximation factor is essentially tight.

cs.DS

Dominance Product and High-Dimensional Closest Pair under $L_\infty$

Given a set $S$ of $n$ points in $\mathbb{R}^d$, the Closest Pair problem is to find a pair of distinct points in $S$ at minimum distance. When $d$ is constant, there are efficient algorithms that solve this problem, and fast approximate solutions for general $d$. However, obtaining an exact solution in very high dimensions seems to be much less understood. We consider the high-dimensional $L_\infty$ Closest Pair problem, where $d=n^r$ for some $r > 0$, and the underlying metric is $L_\infty$. We improve and simplify previous results for $L_\infty$ Closest Pair, showing that it can be solved by a deterministic strongly-polynomial algorithm that runs in $O(DP(n,d)\log n)$ time, and by a randomized algorithm that runs in $O(DP(n,d))$ expected time, where $DP(n,d)$ is the time bound for computing the {\em dominance product} for $n$ points in $\mathbb{R}^d$. That is a matrix $D$, such that $D[i,j] = \bigl| \{k \mid p_i[k] \leq p_j[k]\} \bigr|$; this is the number of coordinates at which $p_j$ dominates $p_i$. For integer coordinates from some interval $[-M, M]$, we obtain an algorithm that runs in $\tilde{O}\left(\min\{Mn^{ω(1,r,1)},\, DP(n,d)\}\right)$ time, where $ω(1,r,1)$ is the exponent of multiplying an $n \times n^r$ matrix by an $n^r \times n$ matrix. We also give slightly better bounds for $DP(n,d)$, by using more recent rectangular matrix multiplication bounds. Computing the dominance product itself is an important task, since it is applied in many algorithms as a major black-box ingredient, such as algorithms for APBP (all pairs bottleneck paths), and variants of APSP (all pairs shortest paths).

cs.DS

Improved Bounds for 3SUM, $k$-SUM, and Linear Degeneracy

Given a set of $n$ real numbers, the 3SUM problem is to decide whether there are three of them that sum to zero. Until a recent breakthrough by Grønlund and Pettie [FOCS'14], a simple $Θ(n^2)$-time deterministic algorithm for this problem was conjectured to be optimal. Over the years many algorithmic problems have been shown to be reducible from the 3SUM problem or its variants, including the more generalized forms of the problem, such as $k$-SUM and $k$-variate linear degeneracy testing ($k$-LDT). The conjectured hardness of these problems have become extremely popular for basing conditional lower bounds for numerous algorithmic problems in P. In this paper, we show that the randomized $4$-linear decision tree complexity of 3SUM is $O(n^{3/2})$, and that the randomized $(2k-2)$-linear decision tree complexity of $k$-SUM and $k$-LDT is $O(n^{k/2})$, for any odd $k\ge 3$. These bounds improve (albeit randomized) the corresponding $O(n^{3/2}\sqrt{\log n})$ and $O(n^{k/2}\sqrt{\log n})$ decision tree bounds obtained by Grønlund and Pettie. Our technique includes a specialized randomized variant of fractional cascading data structure. Additionally, we give another deterministic algorithm for 3SUM that runs in $O(n^2 \log\log n / \log n )$ time. The latter bound matches a recent independent bound by Freund [Algorithmica 2017], but our algorithm is somewhat simpler, due to a better use of word-RAM model.

cs.DS

Coping with Physical Attacks on Random Network Structures

Communication networks are vulnerable to natural disasters, such as earthquakes or floods, as well as to physical attacks, such as an Electromagnetic Pulse (EMP) attack. Such real-world events happen at specific geographical locations and disrupt specific parts of the network. Therefore, the geographical layout of the network determines the impact of such events on the network's physical topology in terms of capacity, connectivity, and flow. Recent works focused on assessing the vulnerability of a deterministic network to such events. In this work, we focus on assessing the vulnerability of (geographical) random networks to such disasters. We consider stochastic graph models in which nodes and links are probabilistically distributed on a plane, and model the disaster event as a circular cut that destroys any node or link within or intersecting the circle. We develop algorithms for assessing the damage of both targeted and non-targeted (random) attacks and determining which attack locations have the expected most disruptive impact on the network. Then, we provide experimental results for assessing the impact of circular disasters to communications networks in the USA, where the network's geographical layout was modeled probabilistically, relying on demographic information only. Our results demonstrates the applicability of our algorithms to real-world scenarios. Our algorithms allows to examine how valuable is public information about the network's geographical area (e.g., demography, topography, economy) to an attacker's destruction assessment capabilities in the case the network's physical topology is hidden or examine the affect of hiding the actual physical location of the fibers on the attack strategy. Thereby, our schemes can be used as a tool for policy makers and engineers to design more robust networks and identifying locations which require additional protection efforts.

eess.SY