arXiv · 2601.17359
Breaking Flat: A Generalised Query Performance Prediction Evaluation Framework
Abstract
The traditional use-case of query performance prediction (QPP) is to identify which queries perform well and which perform poorly for a given ranking model. A more fine-grained and arguably more challenging extension of this task is to determine which ranking models are most effective for a given query. In this work, we generalize the QPP task and its evaluation into three settings: (i) SingleRanker MultiQuery (SRMQ-PP), corresponding to the standard use case; (ii) MultiRanker SingleQuery (MRSQ-PP), which evaluates a QPP model's ability to select the most effective ranker for a query; and (iii) MultiRanker MultiQuery (MRMQ-PP), which considers predictions jointly across all query ranker pairs. Our results show that (a) the relative effectiveness of QPP models varies substantially across tasks (SRMQ-PP vs. MRSQ-PP), and (b) predicting the best ranker for a query is considerably more difficult than predicting the relative difficulty of queries for a given ranker.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Payel Santra, Partha Basuchowdhuri, Debasis Ganguly. 2026-01-24. Breaking Flat: A Generalised Query Performance Prediction Evaluation Framework. https://arxiv.org/abs/2601.17359
Cite the original work for its findings. Save a collection to share your selection of sources.