arXiv · 2603.05750
NERdME: a Named Entity Recognition Dataset for Indexing Research Artifacts in Code Repositories
Abstract
Existing scholarly information extraction (SIE) datasets focus on scientific papers and overlook implementation-level details in code repositories. README files describe datasets, source code, and other implementation-level artifacts, however, their free-form Markdown offers little semantic structure, making automatic information extraction difficult. To address this gap, NERdME is introduced: 200 manually annotated README files with over 10,000 labeled spans and 10 entity types. Baseline results using large language models and fine-tuned transformers show clear differences between paperlevel and implementation-level entities, indicating the value of extending SIE benchmarks with entity types available in README files. A downstream entity-linking experiment was conducted to demonstrate that entities derived from READMEs can support artifact discovery and metadata integration.
Explore related subjects
Keep this discovery
Genet Asefa Gesese, Zongxiong Chen, Shufan Jiang, Mary Ann Tan, Zhaotai Liu, Sonja Schimmler, Harald Sack. 2026-03-05. NERdME: a Named Entity Recognition Dataset for Indexing Research Artifacts in Code Repositories. https://doi.org/10.1145/3774904.3792934
Cite the original work for its findings. Save a collection to share your selection of sources.