arXiv · 2302.14177
Soft-Search: Two Datasets to Study the Identification and Production of Research Software
Abstract
Software is an important tool for scholarly work, but software produced for research is in many cases not easily identifiable or discoverable. A potential first step in linking research and software is software identification. In this paper we present two datasets to study the identification and production of research software. The first dataset contains almost 1000 human labeled annotations of software production from National Science Foundation (NSF) awarded research projects. We use this dataset to train models that predict software production. Our second dataset is created by applying the trained predictive models across the abstracts and project outcomes reports for all NSF funded projects between the years of 2010 and 2023. The result is an inferred dataset of software production for over 150,000 NSF awards. We release the Soft-Search dataset to aid in identifying and understanding research software production: https://github.com/si2-urssi/eager
Explore related subjects
Keep this discovery
Eva Maxfield Brown, Lindsey Schwartz, Richard Lewei Huang, Nicholas Weber. 2023-02-27. Soft-Search: Two Datasets to Study the Identification and Production of Research Software. https://arxiv.org/abs/2302.14177
Cite the original work for its findings. Save a collection to share your selection of sources.