Searcharxiv⌕ Search

arXiv subjects

Martin G. Herold

Publications and source records attributed to Martin G. Herold.

4 recordsLinked to original sources

Can We Break Fine-Grained and NP-Hardness Barriers if We've Seen the Graph Before? The Isomorphic-Priors Model

If we run a heavy-duty computation on prior data, can we avoid repeated computation for similar future inputs? Inspired by this question, we introduce a new computational model for graph problems called algorithms with isomorphic priors. Solving a graph problem $Π$ in this model involves two phases: (i) The preprocessing phase quickly analyzes prior graphs $G_1, ..., G_k$ along with the (previously computed) exact optimal values OPT$(G_i)$. (ii) Subsequently, given a new graph H, a fast query phase must either (a) output the exact solution OPT(H), or (b) correctly report that H is not isomorphic to any $G_i$. Can we avoid computing OPT(H) from scratch when H is isomorphic to some $G_i$? We show that this is the case for a number of problems; for many others, we establish conditional lower bounds. $\textbf{(1)}$ Some NP-hard problems, including Constrained Shortest Path and $\ell_p$-Shortest Path and Constrained Spanning Tree, admit polynomial preprocessing and query times in our model. In contrast, almost all of Karp's 21 NP-complete problems and $(2-\varepsilon)$-approximate $k$-Center, for every fixed $\varepsilon>0$, admit no such algorithms unless Graph Isomorphism (GI) is in P, even with O(1) priors. $\textbf{(2)}$ In contrast to conditional $n^{3-o(1)}$ fine-grained lower bounds, our framework achieves an $O(n^ω)$ query time for Negative Triangle and a near-linear query time for Replacement Path. It also achieves near-linear query time for Maximum Flow. $\textbf{(3)}$ While it remains a major open problem whether infinite-duration games (Paritiy Game, Mean Payoff Game, Energy Game, and Stochastic Game) admit polynomial-time algorithms, they can be easily solved in near-linear time within our model. Our proofs rely on a simple combination of existing tools and are accessible to readers without specialized background.

cs.DS↗

A Broader View on Clustering under Cluster-Aware Norm Objectives

We revisit the $(f,g)$-clustering problem that we introduced in a recent work [SODA'25], and which subsumes fundamental clustering problems such as $k$-Center, $k$-Median, Min-Sum of Radii, and Min-Load $k$-Clustering. This problem assigns each of the $k$ clusters a cost determined by the monotone, symmetric norm $f$ applied to the vector distances in the cluster, and aims at minimizing the norm $g$ applied to the vector of cluster costs. Previously, we focused on certain special cases for which we designed constant-factor approximation algorithms. Our bounds for more general settings left, however, large gaps to the known bounds for the basic problems they capture. In this work, we provide a clearer picture of the approximability of these more general settings. First, we design an $O(\log^2 n)$-approximation algorithm for $(f, L_{1})$-clustering for any $f$. This improves upon our previous $\widetilde{O}(\sqrt{n})$-approximation. Second, we provide an $O(k)$-approximation for the general $(f,g)$-clustering problem, which improves upon our previous $\widetilde{O}(\sqrt{kn})$-approximation algorithm and matches the best-known upper bound for Min-Load $k$-Clustering. We then design an approximation algorithm for $(f,g)$-clustering that interpolates, up to polylog factors, between the best known bounds for $k$-Center, $k$-Median, Min-Sum of Radii, Min-Load $k$-Clustering, (Top, $L_{1}$)-clustering, and $(L_{\infty},g)$-clustering based on a newly defined parameter of $f$ and $g$.

cs.DS↗

Sublinear Data Structures for Nearest Neighbor in Ultra High Dimensions

Geometric data structures have been extensively studied in the regime where the dimension is much smaller than the number of input points. But in many scenarios in Machine Learning, the dimension can be much higher than the number of points and can be so high that the data structure might be unable to read and store all coordinates of the input and query points. Inspired by these scenarios and related studies in feature selection and explainable clustering, we initiate the study of geometric data structures in this ultra-high dimensional regime. Our focus is the {\em approximate nearest neighbor} problem. In this problem, we are given a set of $n$ points $C\subseteq \mathbb{R}^d$ and have to produce a {\em small} data structure that can {\em quickly} answer the following query: given $q\in \mathbb{R}^d$, return a point $c\in C$ that is approximately nearest to $q$. The main question in this paper is: {\em Is there a data structure with sublinear ($o(nd)$) space and sublinear ($o(d)$) query time when $d\gg n$?} In this paper, we answer this question affirmatively. We present $(1+ε)$-approximation data structures with the following guarantees. For $\ell_1$- and $\ell_2$-norm distances: $\tilde O(n \log(d)/\mathrm{poly}(ε))$ space and $\tilde O(n/\mathrm{poly}(ε))$ query time. We show that these space and time bounds are tight up to $\mathrm{poly}{(\log n/ε)}$ factors. For $\ell_p$-norm distances: $\tilde O(n^2 \log(d) (\log\log (n)/ε)^p)$ space and $\tilde O\left(n(\log\log (n)/ε)^p\right)$ query time. Via simple reductions, our data structures imply sublinear-in-$d$ data structures for some other geometric problems; e.g. approximate orthogonal range search, furthest neighbor, and give rise to a sublinear $O(1)$-approximate representation of $k$-median and $k$-means clustering.

cs.DS↗

Clustering to Minimize Cluster-Aware Norm Objectives

We initiate the study of the following general clustering problem. We seek to partition a given set $P$ of data points into $k$ clusters by finding a set $X$ of $k$ centers and assigning each data point to one of the centers. The cost of a cluster, represented by a center $x\in X$, is a monotone, symmetric norm $f$ (inner norm) of the vector of distances of points assigned to $x$. The goal is to minimize a norm $g$ (outer norm) of the vector of cluster costs. This problem, which we call $(f,g)$-Clustering, generalizes many fundamental clustering problems such as $k$-Center, $k$-Median , Min-Sum of Radii, and Min-Load $k$-Clustering . A recent line of research (Chakrabarty, Swamy [STOC'19]) studies norm objectives that are oblivious to the cluster structure such as $k$-Median and $k$-Center. In contrast, our problem models cluster-aware objectives including Min-Sum of Radii and Min-Load $k$-Clustering. Our main results are as follows. First, we design a constant-factor approximation algorithm for $(\textsf{top}_\ell,\mathcal{L}_1)$-Clustering where the inner norm ($\textsf{top}_\ell$) sums over the $\ell$ largest distances. Second, we design a constant-factor approximation\ for $(\mathcal{L}_\infty,\textsf{Ord})$-Clustering where the outer norm is a convex combination of $\textsf{top}_\ell$ norms (ordered weighted norm).

cs.DS↗