Searcharxiv⌕ Search

arXiv · 2610.02385

Fine-Grained Analysis of SIMD-Based Hash Table Implementations

Abstract

In recent years, several variants of classical hash table schemes have been developed by engineers in order to take advantage of the processor's internal parallelism using SIMD instructions, which make it possible to operate on multiple bytes simultaneously. At a small additional memory cost, this enables a significant speedup, making it a data structure that is increasingly popular in practice, when very high performance is required. In this article, we provide a detailed theoretical analysis of the dynamics of such hash tables. From a methodological standpoint, we use and adapt a technique developed by Wormald in the 1990s to study dynamic graphs. This approach, which can be adapted to many variants, enables us to accurately estimate the quantities of interest by capturing the dynamics of the data structure through systems of differential equations. Although complex, we provide an explicit description of the solutions of these systems, which can furthermore be efficiently approximated numerically. Our main results are stated with high probability, which is significantly more precise than average-case analyses, and they match experimental results remarkably well, even for hash tables of moderate size.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Cyril Nicaud, Pablo Rotondo. 2026-10-01. Fine-Grained Analysis of SIMD-Based Hash Table Implementations. https://arxiv.org/abs/2610.02385

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Non-adaptive Bellman-Ford: Yen's improvement is optimal

The Bellman-Ford algorithm for single-source shortest paths repeatedly updates tentative distances in an operation called {relaxing an edge}. In several important applications a {non-adaptive} (oblivious) implementation is preferred, which means fixing the entire sequence of relaxations upfront, independently of the edge-weights. The original implementation of the algorithm performs, in a dense graph on $n$ vertices, $(1+o(1))n^3 $ relaxations. An improvement by Yen from 1970 reduces the number of relaxations by a factor of two. We show that no further constant-factor improvements are possible, and every {non-adaptive deterministic} algorithm based on relaxations must perform $(\frac{1}{2} - o(1))n^3$ steps. This improves an earlier lower bound of Eppstein of $(\frac{1}{6} - o(1))n^3$. Given that a {non-adaptive randomized} variant of Bellman-Ford with at most $(\frac{1}{3} + o(1))n^3$ relaxations (with high probability) is known, our result implies a strict separation between deterministic and randomized strategies, answering an open question of Eppstein. We also address the complexity of finding {short} relaxation sequences for a given input graph on $n$ vertices, answering a question of Eppstein. We show that the problem is co-NP-hard, and moreover essentially inapproximable: While an $n$-approximation is easily obtained, for every $ε> 0$, no polynomial-time $n^{1-ε}$-approximation exists, unless P = NP. We further show that {deciding} whether a given relaxation sequence is valid is co-NP-complete, even when the input is the complete graph.

cs.DS↗

Exponential Quantum Space Advantage for Approximating Max-$k$SAT in the Streaming Setting

In this paper, we give a one-pass quantum streaming algorithm for Max-$k$SAT that uses $\operatorname{polylog}(n)$ space and achieves a $0.7426$-approximation on instances with $n$ variables and $\operatorname{poly}(n)$ clauses. In contrast, prior work by Chou, Golovnev, and Velusamy (FOCS 2020) implies that achieving an approximation ratio better than $\sqrt{2}/2 \approx 0.7071$ for Max-$k$SAT requires $Ω(\sqrt{n})$ space for any classical streaming algorithm. Therefore, it yields an exponential quantum space advantage for Max-$k$SAT in the streaming setting. Combining with the known results, it gives a complete classification of quantum space advantages for all Boolean Max-2CSPs.

cs.DS↗

Solving Stackelberg Vertex Cover on trees using split and join

The Stackelberg Vertex Cover problem is a bilevel optimization problem with two players on a graph G = ($F \cup P$, E) where each vertex from F has a weight and the first player selects a price for each vertex in P . Afterwards, the second player finds a minimum weight vertex cover X and the first player receives the set price for each vertex from $X \cap P$ . The goal is to maximize the revenue of the first player. This problem was recently shown to be NP-complete for bipartite graphs while being solvable in linear time on paths. We present four new algorithms for solving Stackelberg Vertex Cover on certain kinds of graphs: (1) a pseudo-polynomial algorithm working on general trees when all weights are integer with a runtime linear in the number of vertices and cubic in the maximum weight (2) a generalization of (1) for bipartite graphs with integer weights and a tree decomposition that is FPT in the maximum weight and the treewidth, (3) a strongly polynomial algorithm for rooted trees having the property that the least common ancestor of any two vertices from P is again in P (this case includes paths); and (4) an FPT-algorithm for trees, where the parameter is the maximum number P-vertices $v_i$ that an F-vertex u can reach while using no other P -vertices. These algorithms are based on a lemma that allows us to split instances at a vertex u into multiple sub-instances, which follows from LP duality and integrality of the vertex cover LP on bipartite graphs. The lemma requires that the minimum vertex covers of the sub-instances agree on u (either all include u or all don't). For this we introduce the concept of commitments. We show that the Stackelberg Vertex Cover problem with commitments is weakly NP-complete. An open question is the non-bipartite case as there is an explicit counterexample showing that the split-and-join technique does not work.

cs.DS↗