Searcharxiv⌕ Search

arXiv · 2609.39557

Consensus for Compressed Static Functions

Abstract

The Consensus technique marked a breakthrough in the construction of minimal perfect hash functions (MPHFs), reaching a linear tradeoff between construction time and space overhead relative to the optimum. Consensus provides a clever scheme to search for and encode seeds of tasks in random data structures. We apply Consensus to the related field of compressed static functions (CSFs). These data structures store a function $f: S \to Σ$ such that querying a key $x \in S$ returns $f(x)$ and querying $x \not \in S$ returns an arbitrary value. CSFs do not need to store the keys $S$ and only need space close to the zeroth-order empirical entropy of the multiset of values. Often, some values are much more common than others. In these cases, CSFs can use less space than their non-compressed counterparts. CSFs are a useful building block, for example in database design and bioinformatics. We introduce Consensus-CSF, which can reach arbitrarily close to the empirical entropy $n H_0$, with a construction time of $n \exp(\tilde{\cal{O}} (\sqrt{1 / δ}))$ for space usage of $n H_0 (1 + δ)$ when assuming some parameters of the value distribution to be constants. This tradeoff beats previously implemented approaches that can only reach some fixed threshold above the entropy lower bound. We enable Consensus in the setting of CSFs, which is less structured than MPHFs, with the introduction of task insertions. Our approach randomly distributes the keys into one-bit Consensus tasks and then strategically inserts additional tasks in places where the construction would get stuck otherwise. We provide an implemented version of our algorithm which reaches the same order of magnitude in space overhead as competitors but is not competitive in practice. Beyond these results, we present a new way to think and reason about Consensus, which may also be applied to other problems.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Dominik Rosch, Jonatan Ziegler. 2026-09-30. Consensus for Compressed Static Functions. https://arxiv.org/abs/2609.39557

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Online Matching and Contention Resolution for Edge Arrivals with Vanishing Probabilities

We study the performance of sequential contention resolution and matching algorithms on random graphs with vanishing edge probabilities. When the edges of the graph are processed in an adversarially-chosen order, we derive a new OCRS that is $0.382$-selectable, attaining the "independence benchmark" from the literature under the vanishing edge probabilities assumption. Complementary to this positive result, we show that no OCRS can be more than $0.390$-selectable, significantly improving upon the upper bound of $0.428$ from the literature. We also derive negative results that are specialized to bipartite graphs or subfamilies of OCRSs. Meanwhile, when the edges of the graph are processed in a uniformly random order, we show that the simple greedy contention resolution scheme which accepts all active and feasible edges is $1/2$-selectable. This result is tight due to a known upper bound. We then show that when the algorithm can choose the processing order, a slight tweak to the random order---give each vertex a random priority and process edges in lexicographic order---results in a strictly better contention resolution scheme that is $1-\ln(2-1/e)\approx0.510$-selectable. Moreover, we show that this bound is tight over any sequential contention resolution scheme, even one which may adaptively choose the order in which it processes edges. This provides a separation from the $0.544$ upper bound for offline contention resolution implied by the classic result of Karp and Sipser. Our positive results also apply to online matching on $1$-uniform random graphs with vanishing (non-identical) edge probabilities, extending and unifying some results from the random graphs literature.

cs.DS↗

Counting large patterns in degenerate graphs

The problem of subgraph counting asks for the number of occurrences of a pattern graph $H$ as a subgraph of a host graph $G$ and is known to be computationally challenging: it is $\#W[1]$-hard even when $H$ is restricted to simple structures such as cliques or paths. Curticapean and Marx (FOCS'14) show that if the graph $H$ has vertex cover number $τ$, subgraph counting has time complexity $O(|H|^{2^{O(τ)}} |G|^{τ+ O(1)})$. This raises the question of whether this upper bound can be improved for input graphs $G$ from a restricted family of graphs. Earlier work by Eppstein~(IPL'94) shows that this is indeed possible, by proving that when $G$ is a $d$-degenerate graph and $H$ is a biclique of arbitrary size, subgraph counting has time complexity $O(d 3^{d/3} |G|)$. We show that if the input is restricted to $d$-degenerate graphs, the upper bound of Curticapean and Marx can be improved for a family of graphs $H$ that includes all bicliques and satisfies a property we call $(c,d)$-locatable. Importantly, our algorithm's running time only has a polynomial dependence on the size of~$H$. A key feature of $(c,d)$-locatable graphs $H$ is that they admit a vertex cover of size at most $cd$. We further characterize $(1,d)$-locatable graphs, for which our algorithms achieve a linear running time dependence on $|G|$, and we establish a lower bound showing that counting graphs which are barely not $(1,d)$-locatable is already $\#\text{W}[1]$-hard. We note that the restriction to $d$-degenerate graphs has been a fruitful line of research leading to two very general results (FOCS'21, SODA'25) and this creates the impression that we largely understand the complexity of counting substructures in degenerate graphs. However, all aforementioned results have an exponential dependency on the size of the pattern graph $H$.

cs.DS↗

Budget-Independent Influence Maximization in Nearly Linear Time

Influence maximization asks for $k$ seed vertices that maximize the expected spread of a diffusion process in a network. Standard near-optimal-time algorithms based on reverse-reachable sampling achieve a $(1-1/e-\varepsilon)$ approximation, but their expected running-time bounds grow linearly with the seed budget $k$. We remove this multiplicative dependence: for the independent cascade model, our algorithm succeeds with probability at least $1-δ$ in $O((m+n)\varepsilon^{-3}\log(2n/δ))$ expected time. The result extends to triggering models with explicitly charged local sampling costs. We reserve $O(\varepsilon k)$ seed positions for cost-weighted random vertices, allowing reverse-reachable searches to stop as soon as they encounter a reserved seed. An independent sample-count estimation phase uses a statistic that also controls the expected search cost. Matching these quantities eliminates the multiplicative dependence on $k$ while preserving the approximation guarantee.

cs.DS↗