arXiv · 2610.03130
Benchmarking Literature Retrieval for a Model Organism: A Dictyostelium Case Study
Abstract
Biological literature retrieval systems are often developed and evaluated using broad biomedical corpora and general-purpose search tasks. However, many curated knowledge bases operate in narrower model-organism domains, where the literature is sparse and terminology is organism-specific. We introduce a retrieval benchmark from dictyBase for Dictyostelium, a model organism in cell and developmental biology. The benchmark consists of curator-generated biological queries linked to PubMed-indexed articles, together with structured gene annotations. Using this benchmark, we study three factors in niche biological retrieval: cross-encoder reranking, gene-aware query expansion, and abstract-only versus full-text retrieval. We report that reranking and gene-aware query expansion improve retrieval selectively: reranking is most useful when the model is well suited to biological evidence matching, whereas curated annotations help clarify compact biological queries by reducing vocabulary mismatch. Full-text chunks substantially improve retrieval when abstracts omit supporting evidence, increasing both candidate recall and top-rank performance, although these cases are harder than queries supported by abstracts. Data and code are publicly available at https://github.com/fulaibaowang/dictycite, and the benchmark dataset is additionally archived on Zenodo.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yun Wang, Gad Shaulsky, Tomaž Curk, Blaž Zupan. 2026-10-02. Benchmarking Literature Retrieval for a Model Organism: A Dictyostelium Case Study. https://arxiv.org/abs/2610.03130
Cite the original work for its findings. Save a collection to share your selection of sources.