arXiv · 2609.19071
Benchmarking Large Language Models for Biomedical Relation Extraction
Abstract
Extracting SNP-phenotype associations from biomedical literature is vital but challenging. We benchmarked diverse NLP models, including MLMs, hybrid architectures, and state-of-the-art LLMs (Gemini 2.0, OpenAI O-series, Qwen, Mistral), on the SNPPhenA corpus across three tasks: sentence-level, abstract-level, and association strength classification. OpenAI O1 achieved state-of-the-art (SOTA) results using few-shot learning for non-finetuned sentence-level classification (F1 0.89) and established a new SOTA for abstract-level classification (F1 0.82). Association strength classification proved difficult, though fine-tuned Gemini 2.0 Pro performed best (F1 0.60) in the first LLM evaluation of this task. Proprietary LLMs, especially in few-shot (O1) or fine-tuned (Gemini 2.0 Pro) settings, significantly outperformed other models. These findings confirm the power of modern LLMs for genomic knowledge extraction.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Claudiu Creanga, Teodor Marchitan, Liviu P. Dinu. 2026-07-23. Benchmarking Large Language Models for Biomedical Relation Extraction. https://arxiv.org/abs/2609.19071
Cite the original work for its findings. Save a collection to share your selection of sources.