Searcharxiv⌕ Search

arXiv · 2610.09412

Finding the Right Balance: Relevance and Diversity in LLM Retrieval

Abstract

Retrieval diversification is widely available in retrieval-augmented generation (RAG) frameworks, yet prior studies disagree on whether it improves retrieval and answer quality. We show that its effectiveness varies primarily with candidate-pool redundancy, in a pattern consistent with the number of distinct evidence pieces a query requires. Using controlled near-duplicate injection and production-style overlapping chunking, we find that diversification harms relevance, evidence coverage and answer quality on clean pools, but becomes beneficial on multi-evidence tasks when redundancy causes nearest-neighbor retrieval to select repeated passages. We therefore introduce a query-adaptive rule that diversifies only when the effective number of distinct documents in the nearest-neighbor top-$k$ selection falls below the query's evidence requirement. Computed from existing embeddings, the rule captures most of the achievable gain, transfers across datasets and encoders and automatically reduces to nearest-neighbor retrieval for single-evidence queries. We also introduce RNG-Score, a geometric reranker with an exact nearest-neighbor fallback whose margin indicates duplicate structure. Overall, we conclude that diversification should be used selectively, based on observable redundancy and evidence requirements.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Guillaume Brouillette, Faustin Kagabo, Usef Faghihi, Nadia Ghazzali. 2026-10-07. Finding the Right Balance: Relevance and Diversity in LLM Retrieval. https://arxiv.org/abs/2610.09412

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

H2CE: Modeling Geo-Semantic Interactions for POI Reranking with Heterogeneous Two-Stage Cross-Encoders

Point-of-Interest (POI) reranking in local search must model query-conditioned tradeoffs among lexical semantics, geospatial proximity, and numerical quality signals such as rating and review count, while remaining practical under real-time serving constraints. A close POI may only partially satisfy the query intent, while a farther one may offer stronger semantic and quality evidence. We present H2CE, a Heterogeneous Two-stage Cross-Encoder for latency-bounded POI reranking. H2CE represents numerical attributes in two complementary ways: bucketized natural-language descriptors are inserted into the cross-encoder input to support semantic--numeric attention, while exact scalar values are processed by dedicated MLPs to preserve magnitude information. The resulting semantic and numerical embeddings are fused through latent-space aggregation, enabling nonlinear interactions beyond scalar weighted sums. H2CE then applies a two-stage architecture: Stage 1 scores all candidates pointwise for scalable filtering, and Stage 2 performs head-to-head pairwise comparison among the top-$K$ candidates with Copeland aggregation, making fine-grained relative tradeoffs explicit while reducing pairwise cost from O(N^2) to O(N+K(K-1)). On a 5,743-query local search test set, H2CE achieves 67.48% NDCG@5, improving over XGBoost LTR by +22.82% absolute and over a zero-shot LLM reranker by +35.89%. The pairwise stage adds +1.98% NDCG@5 over the pointwise model alone. Ablations confirm the value of numerical features, latent aggregation, top-K pairwise reranking, and aligned training.

cs.IR↗

Generative Spatiotemporal Intent Sequence Recommendation via Implicit Reasoning in Amap

Real-world user behavior rarely consists of isolated actions; instead, it often forms intent flows governed by spatiotemporal dependencies. To provide integrated service recommendations, we focus on the task of Generative Spatiotemporal Intent Sequence Recommendation (GSISR), which aims to generate intent sequences that are logically coherent and physically executable within complex spatiotemporal contexts. While LLMs offer strong reasoning potential for GSISR, direct industrial deployment is limited by high inference latency and context-mismatched or physically infeasible plans. To address these challenges, we propose a generative framework, GPlan, that internalizes LLM reasoning into lightweight models through two components. First, to enable reasoning under strict latency constraints, we introduce Progressive Implicit CoT Distillation, which compresses explicit reasoning processes into reserved latent tokens, allowing small models to inherit complex planning logic without generating long reasoning text. Second, to address the disconnect between general knowledge and real-world constraints, we design Spatiotemporal Counterfactual DPO. By aligning the model with counterfactual context-plan pairs, we improve sensitivity to spatiotemporal context and reduce context-mismatched plans. Offline experiments and online A/B testing demonstrate that our approach improves sequence coherence and context responsiveness. Our implementation and the anonymized GSISR dataset are available at https://github.com/alibaba/GPlan.

cs.IR↗

Hypergraph-Enhanced Dual Convolutional Network for Bundle Recommendation

Bundle recommendation ranks sets of related items rather than isolated items. Its central challenge is to connect user preferences, item interactions, and bundle composition without losing the signals needed to rank bundles. We propose Hypergraph-Enhanced Dual Convolutional Neural Network (HED), which constructs a complete hypergraph containing user--bundle, user--item, and bundle--item interactions together with intra-user and intra-bundle relations. HED couples complete-hypergraph propagation with a user--bundle branch, allowing item-aware higher-order context to inform ranking while preserving recommendation-specific signals. On NetEase, HED-128 improves over the strongest baseline by 5.04--6.97% across the six reported metrics; on Youshu, HED-64 improves by 1.87--4.56%. Ablation results support the contributions of both the user--bundle branch and intra-type relations, and sensitivity analyses identify stable operating ranges for the main hyperparameters. We further quantify the computational trade-off of the complete hypergraph, including its memory cost. The evidence supports HED on the two evaluated bundle-recommendation datasets while making its resource limitations explicit. Code and datasets will be made available upon publication.

cs.IR↗