arXiv · 2506.23191
Impact of Shallow vs. Deep Relevance Judgments on BERT-based Reranking Models
Abstract
This paper investigates the impact of shallow versus deep relevance judgments on the performance of BERT-based reranking models in neural Information Retrieval. Shallow-judged datasets, characterized by numerous queries each with few relevance judgments, and deep-judged datasets, involving fewer queries with extensive relevance judgments, are compared. The research assesses how these datasets affect the performance of BERT-based reranking models trained on them. The experiments are run on the MS MARCO and LongEval collections. Results indicate that shallow-judged datasets generally enhance generalization and effectiveness of reranking models due to a broader range of available contexts. The disadvantage of the deep-judged datasets might be mitigated by a larger number of negative training examples.
Explore related subjects
Keep this discovery
Gabriel Iturra-Bocaz, Danny Vo, Petra Galuscakova. 2025-06-29. Impact of Shallow vs. Deep Relevance Judgments on BERT-based Reranking Models. https://doi.org/10.1145/3731120.3744602
Cite the original work for its findings. Save a collection to share your selection of sources.