Searcharxiv⌕ Search

arXiv · 2610.10135

Attention via Black-Box Vector Search

Abstract

Sparse attention mechanisms estimate attention over $n$ tokens using a small subset of keys. Many existing approaches use maximum inner product search (MIPS) to retrieve the heaviest keys, which motivates the following question: given black-box access to a MIPS oracle, how many keys must be retrieved to output an $\varepsilon$-accurate attention estimate? We answer this question by unifying prior approaches through the framework of priority sampling. With a single MIPS index, we show that $Θ(\sqrt{n}/\varepsilon)$ retrieved keys are both sufficient and necessary. With $Θ(\log n)$ indices, we give an algorithm that retrieves only $O(\log n+1/\varepsilon^2)$ keys and prove that this is near-optimal. More generally, we design algorithms that establish a smooth tradeoff between the number of MIPS indices and number of retrieved keys. We then show that if we allow augmentation of keys and queries, we can bypass the above lower bounds: there exists a simple priority-sampling estimator using a single MIPS index and $O(1/\varepsilon^2)$ retrieved keys. When integrated into LLM inference, our algorithms outperform top-$k$ and sampling approaches used in prior work and yield attention approximation that scales favorably to long contexts.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Stepan Zharkov, Krish Singal, Ashwin Padaki, Alexandr Andoni. 2026-10-07. Attention via Black-Box Vector Search. https://arxiv.org/abs/2610.10135

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Almost Optimal Constant-Round Approximation of Dominating Set in Graph Classes with Excluded Minors

For every fixed proper minor-closed class $\mathscr C$ and every $ε>0$, we give a deterministic LOCAL algorithm that returns a dominating set of size at most $(2a(\mathscr C)+1+ε)γ_f(G)$ on every $G\in\mathscr C$. Here $a(\mathscr C)$ is the supremum edge-to-vertex ratio in $\mathscr C$, and $γ_f(G)$ is the fractional domination number. The class also admits a deterministic $(1+ε)$-approximation for fractional dominating set and a randomized algorithm that always returns a dominating set and has expected size at most $(1+ε)γ(G)$. In each case, the number of rounds depends only on $\mathscr C$ and $ε$. None of these algorithms requires the number of vertices or the maximum degree as part of the input. For planar graphs, this gives the deterministic guarantee $(7+ε)γ_f(G)$. Together with the lower bound of Hilke, Lenzen and Suomela, it determines the infimum of the deterministic constant-round approximation ratios for planar minimum dominating set as $7$, settling a question that had remained open since their work. The corresponding infima, measured against the integral optimum, are $7$ for graphs of Euler genus at most any fixed $g\ge0$, $2t-3$ for $K_t$-minor-free graphs with $3\le t\le9$, and $2r+1$ for graphs of treewidth or pathwidth at most any fixed $r\ge1$. We also prove that, for every integer $r\ge1$, no deterministic constant-round LOCAL algorithm achieves an approximation ratio below $2r+1$ on the $r$-th powers of paths, even when every vertex knows the number of vertices. This gives a new proof that the limiting constants are optimal for planar graphs, graphs of bounded treewidth or pathwidth, and $K_t$-minor-free graphs with $3\le t\le9$. For triangle-free planar graphs, the corresponding infimum is $5$.

cs.DS↗

Truly Sub-$3^n$ Min-Sum Subset Convolution and Join Ordering

We present a deterministic reduction from min-sum subset convolution to min-plus matrix product. We show that if the min-plus product of two $D\times D$ matrices with $β$-bit integer entries can be computed in $D^{3-δ}\operatorname{poly}(β,\log D)$ time for a fixed rational $0<δ<1$, then min-sum subset convolution on an $n$-element universe can be solved in $(2+2^{-δ})^n 2^{O(\sqrt n\log(n+1))}\operatorname{poly}(n,β)$ time. Instantiating this reduction with the recent breakthrough on subcubic min-plus matrix product by Alman and Vassilevska Williams gives a Las Vegas algorithm with expected running time $O^*(2.9987^n)$ and a deterministic algorithm with running time $O^*(2.9997^n)$, strictly breaking the longstanding $3^n$ computational barrier. Notably, these speedups translate directly to database query optimization, yielding the same expected and deterministic running-time bounds for join ordering under the $C_{\mathrm{out}}$ cost function.

cs.DS↗

Tight Bounds for Equivalence Testing with Non-Adaptive Conditional Samples

We study distribution testing with access to non-adaptive conditional samples. Specifically, we give tight bounds for equivalence testing, determining whether two unknown distributions are equal to or $\varepsilon$-far from each other in total variation distance. Our algorithm and lower bound show that $\tilde Θ\left(\frac{\log n}{\varepsilon^2}\right)$ queries are necessary and sufficient for this problem. These results demonstrate that the complexity of uniformity, identity, and equivalence testing with non-adaptive conditional samples are all $\tilde Θ(\log n)$.

cs.DS↗