Searcharxiv⌕ Search

arXiv · 2609.40249

Dynamic Time Warping in the Low-Distance Regime

Abstract

Dynamic Time Warping (DTW) is a classical similarity measure for strings and time series that allows local stretching. Given non-empty strings $S,T$ over an alphabet $Σ$ and a cost function $δ:Σ^2\to\mathbb{R}_{\ge0}$, $DTW_δ(S,T)$ is the minimum total cost of equal-length expansions of $S$ and $T$ obtained by duplicating characters. For strings of length at most $n$, DTW is computable in $O(n^2)$ time, and this is conditionally optimal under the Orthogonal Vectors Hypothesis (OVH). We study the low-distance regime, where an integer $k$ upper-bounds $DTW_δ(S,T)$, assuming $δ(a,a)=0$ and $δ(a,b)\ge1$ for $a\ne b$. For several classical similarity measures, this regime admits $O(n+\operatorname{poly}(k))$ algorithms, whereas for DTW with metric costs the best known bound is $O(nk)$. We show that this dependence is essentially optimal: assuming OVH, computing DTW requires $n^{1-o(1)}k$ time even for the discrete mismatch-cost function, which assigns cost $1$ to every mismatch. The lower bound applies to the whole spectrum of thresholds $k$ between constant and linear in $n$. Our reduction from Orthogonal Vectors encodes vector coordinates in the lengths of equal-character runs. The resulting instances are very structured: collapsing runs to single characters reveals long substrings with short periods. We complement the lower bound with a $\tilde O(n+\operatorname{poly}(k))$-time algorithm whenever, after collapsing runs in the inputs, every substring with period $O(k)$ has length $\operatorname{poly}(k)$. Finally, we extend this lower bound to DTW pattern matching, which asks whether any non-empty substring of a length-$n$ text has DTW distance at most $k$ from a length-$m$ pattern. We prove that the classic $O(nm)$-time dynamic-programming algorithm is near-optimal under OVH, even when $k=O(\log n)$.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Itai Boneh, Shay Golan, Tomasz Kociumaka. 2026-09-30. Dynamic Time Warping in the Low-Distance Regime. https://arxiv.org/abs/2609.40249

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Online Matching and Contention Resolution for Edge Arrivals with Vanishing Probabilities

We study the performance of sequential contention resolution and matching algorithms on random graphs with vanishing edge probabilities. When the edges of the graph are processed in an adversarially-chosen order, we derive a new OCRS that is $0.382$-selectable, attaining the "independence benchmark" from the literature under the vanishing edge probabilities assumption. Complementary to this positive result, we show that no OCRS can be more than $0.390$-selectable, significantly improving upon the upper bound of $0.428$ from the literature. We also derive negative results that are specialized to bipartite graphs or subfamilies of OCRSs. Meanwhile, when the edges of the graph are processed in a uniformly random order, we show that the simple greedy contention resolution scheme which accepts all active and feasible edges is $1/2$-selectable. This result is tight due to a known upper bound. We then show that when the algorithm can choose the processing order, a slight tweak to the random order---give each vertex a random priority and process edges in lexicographic order---results in a strictly better contention resolution scheme that is $1-\ln(2-1/e)\approx0.510$-selectable. Moreover, we show that this bound is tight over any sequential contention resolution scheme, even one which may adaptively choose the order in which it processes edges. This provides a separation from the $0.544$ upper bound for offline contention resolution implied by the classic result of Karp and Sipser. Our positive results also apply to online matching on $1$-uniform random graphs with vanishing (non-identical) edge probabilities, extending and unifying some results from the random graphs literature.

cs.DS↗

Counting large patterns in degenerate graphs

The problem of subgraph counting asks for the number of occurrences of a pattern graph $H$ as a subgraph of a host graph $G$ and is known to be computationally challenging: it is $\#W[1]$-hard even when $H$ is restricted to simple structures such as cliques or paths. Curticapean and Marx (FOCS'14) show that if the graph $H$ has vertex cover number $τ$, subgraph counting has time complexity $O(|H|^{2^{O(τ)}} |G|^{τ+ O(1)})$. This raises the question of whether this upper bound can be improved for input graphs $G$ from a restricted family of graphs. Earlier work by Eppstein~(IPL'94) shows that this is indeed possible, by proving that when $G$ is a $d$-degenerate graph and $H$ is a biclique of arbitrary size, subgraph counting has time complexity $O(d 3^{d/3} |G|)$. We show that if the input is restricted to $d$-degenerate graphs, the upper bound of Curticapean and Marx can be improved for a family of graphs $H$ that includes all bicliques and satisfies a property we call $(c,d)$-locatable. Importantly, our algorithm's running time only has a polynomial dependence on the size of~$H$. A key feature of $(c,d)$-locatable graphs $H$ is that they admit a vertex cover of size at most $cd$. We further characterize $(1,d)$-locatable graphs, for which our algorithms achieve a linear running time dependence on $|G|$, and we establish a lower bound showing that counting graphs which are barely not $(1,d)$-locatable is already $\#\text{W}[1]$-hard. We note that the restriction to $d$-degenerate graphs has been a fruitful line of research leading to two very general results (FOCS'21, SODA'25) and this creates the impression that we largely understand the complexity of counting substructures in degenerate graphs. However, all aforementioned results have an exponential dependency on the size of the pattern graph $H$.

cs.DS↗

Budget-Independent Influence Maximization in Nearly Linear Time

Influence maximization asks for $k$ seed vertices that maximize the expected spread of a diffusion process in a network. Standard near-optimal-time algorithms based on reverse-reachable sampling achieve a $(1-1/e-\varepsilon)$ approximation, but their expected running-time bounds grow linearly with the seed budget $k$. We remove this multiplicative dependence: for the independent cascade model, our algorithm succeeds with probability at least $1-δ$ in $O((m+n)\varepsilon^{-3}\log(2n/δ))$ expected time. The result extends to triggering models with explicitly charged local sampling costs. We reserve $O(\varepsilon k)$ seed positions for cost-weighted random vertices, allowing reverse-reachable searches to stop as soon as they encounter a reserved seed. An independent sample-count estimation phase uses a statistic that also controls the expected search cost. Matching these quantities eliminates the multiplicative dependence on $k$ while preserving the approximation guarantee.

cs.DS↗