Searcharxiv⌕ Search

arXiv · 2610.02572

Adaptive Sparsity Optimization with Learnable Soft Top-K and Per-Term Thresholding for Efficient Retrieval

Abstract

Recent work on neural sparse retrieval has demonstrated strong relevance by leveraging Large Language Models (LLMs) for semantic term expansion. However, learned models paired with previous sparsification techniques still yield overly long document and query vectors partly due to a large LLM vocabulary, imposing a serious challenge to retrieval time and space efficiency. This paper proposes a scheme for optimizing model sparsity through a synergy of adaptive strategies, including learnable soft top-K, per-term thresholding, and FLOPs regularization to increase the sparsity of query and document vectors. Experimental results with Lion-SP model on the MS MARCO and BEIR datasets demonstrate that the proposed scheme can outperform the baselines by significantly reducing the average query and document lengths. Our scheme can achieve much shorter retrieval latency and lower storage cost while maintaining highly competitive relevance.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wentai Xie, Parker Carlson, Shanxiu He, Tao Yang. 2026-10-01. Adaptive Sparsity Optimization with Learnable Soft Top-K and Per-Term Thresholding for Efficient Retrieval. https://doi.org/10.1145/3805712.3809625

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

More Efficient LLM Reranking with Whole-Pool, Setwise, Long-Context Language Models

LLM-based re-rankers produce a rankings through repeated local comparisons (listwise, pairwise or pointwise), requiring many sequential model calls. We study how long-context LLMs can drastically reduce this computation when the entire retrieved candidate pool fits within the context window. We introduce Whole-Pool Setwise re-ranking, where each comparison ranks all the entire candidate pool, and propose DualEnd, which jointly selects the candidates predicted to be most and least relevant. By filling the ranking from both ends, DualEnd constructs a complete ranking of 100 candidates in 50 LLM comparisons. Experiments with nine open-weight LLMs on TREC DL19 and DL20 show that this requires 59.4% fewer comparisons than previous top-oriented windowed Setwise with heapsort and 88.8% fewer than top-oriented windowed Setwise with bubblesort, even though those baselines target only the top-10 rankings while DualEnd targets the full ranking. DualEnd's nDCG@100 is within 0.008 of single-end whole-pool top-oriented approach, while approximately halving its token consumption and ranking time. Across six BEIR datasets, DualEnd reduces mean token consumption and ranking time by 49.4% and 50.8%, respectively, relative to single-end whole-pool top-oriented approach. These results demonstrate that DualEnd Setwise enables complete re-ranking with substantially fewer LLM comparisons and competitive effectiveness across several backbones.

cs.IR↗

Min-Cost Flow Routing for Evidence Assembly in Long Multimodal Documents

Answering questions about long multimodal documents requires distributing a fixed evidence budget across relevant facets in text, tables, figures, and slides while avoiding near-duplicates. We present \flowreader, which formulates evidence selection as a single minimum-cost flow problem with capacity limits over a multimodal content graph. Spectral decomposition identifies latent aspects of query-relevant content and allocates the budget among them in proportion to their spectral energy. These capacity limits enforce aspect coverage during routing without requiring a language-model planning call. Query-conditioned costs prioritize chains of relevant, mutually consistent evidence. Decomposing the optimal flow produces short evidence chains, which a vision-language model reads in parallel and a reasoner reconciles. On VisDoMBench with Qwen3-VL-32B, \flowreader\ achieves the highest macro accuracy ($68.9$), surpassing the strongest prior system by $2.7$ points, leading on three of five subsets and attaining the highest worst-subset accuracy. It uses a measured $17.5$ content nodes per query and maintains its lead at $12.9$. Ablation studies with a fixed graph, scorer, reader, and judge show that cost design drives accuracy, capacity limits preserve it while using about three-quarters of the reader tokens required by shortest-path routing without these limits on the same network, and spectral aspects align with LLM-generated sub-questions without a planning call.

cs.IR↗

Learning Query Encoders Can Be Hard Even When Vector Retrieval Is Geometrically Easy

Efficient vector retrieval requires both a corpus geometry that supports retrieving the right documents through vector similarity, and a query encoder that can embed queries near their desired documents in the embedding space. Recent work has studied geometric capacity through the lens of the minimum embedding dimension needed to realize all top-$k$ answer sets of $n$ documents. We study a different notion of geometric capacity--the maximum recall achievable for a frozen document index--and explore whether learned query encoders can reach this ceiling. On several real-world retrieval benchmarks, we show that retrieval quality of single-vector query encoders often lies far below what the document indices can support. Motivated by this observation, we give theoretical evidence that learning query encoders can be computationally hard. In particular, we construct a retrieval task that (1) admits a query encoder with perfect recall which is representable by a small one-hidden-layer ReLU network, but (2) any statistical-query learner (a class capturing learners that access training data through aggregate statistics) provably requires exponentially many statistical queries to achieve non-trivial recall advantage over the random baseline $k/n$. Taken together, our results suggest substantial unrealized geometric capacity in retrieval benchmarks and establish query encoder learnability as a possible barrier in embedding-based retrieval.

cs.IR↗