arXiv · 2606.13051
AAbAAC: An Annotated Corpus for Autoimmunity Information Extraction
Abstract
Despite advances in information extraction driven by deep learning and large language models, performance gaps remain in highly specialized biomedical fields, where domainspecific complexity poses challenges for generalist models. In this work, we focus on the domain of autoimmunity, where the main entities of interest are autoimmune diseases, autoantibodies (i.e., molecules that may mark or cause these diseases), their molecular targets, their location in the body, and their associated clinical signs. Herein, we present AAbAAC (AutoAntibodies and Autoimmunity Annotated Corpus), a corpus of 115 abstracts selected from PubMed, where we manually annotated entities and their relationships. First, AAbAAC was used to evaluate several methods on the task of named entity recognition (NER), and secondly, to fine-tune NER models. Our study demonstrates the utility of AAbAAC for information extraction in the domain of autoimmunity, showing expected improvement in NER performance after finetuning. This illustrates the value of small-scale annotation efforts for specialized domains and contributes to the computational study of autoimmunity. The AAbAAC corpus is available at https://github.com/f-maury/AAbAAC.
Explore related subjects
Keep this discovery
Fabien Maury, Solène Grosdidier, Maud de Dieuleveult, Adrien Coulet. 2026-06-11. AAbAAC: An Annotated Corpus for Autoimmunity Information Extraction. https://arxiv.org/abs/2606.13051
Cite the original work for its findings. Save a collection to share your selection of sources.