Searcharxiv⌕ Search

arXiv · 2610.03443

Still funded, no longer counted: how NIH's 2025 award reviews changed what the government counts as minority health research

Abstract

Funders know their portfolios through software that classifies award text. The US National Institutes of Health (NIH) reports its spending in more than 300 categories mined this way, and work on classification and indicators treats the text as the applicant's to write. In 2025 NIH required "DEI language" removed from awards not supporting DEI activities, and so policed the words it also counts. We followed population names through 37,790 continuing awards, checking them against practice records. Names were informative: in new awards with trial baselines, a title naming Black populations predicted a 60-percentage-point higher enrolled share. Text features predicted which names survived, and practice records added little. NIH's minority health category followed the names: of continuing projects carrying it, 99.4% with unchanged summaries kept it, against 15.7% of those whose summaries no longer named a racial or ethnic population, while the projects kept their funding. Blinded reviewers judged 42 of a random 50 such losses to have met its definition in FY2024. Comparing both summaries of 70 name losses, they found aims concerning the population recast in 49. Read alone, 63 FY2025 summaries no longer met the definition. Awards naming sexual and gender minorities were recorded as terminated 37.8 points more often after adjustment for listed terms, institute and activity. We call the mechanism word targeting, its boundary population targeting, and its product uncounted science: funded research the count no longer records. The count followed the edited text, and the edited record cannot say whether the research changed with it.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Fangfang Xie, Jingwen Zhang, Haining Wang. 2026-10-02. Still funded, no longer counted: how NIH's 2025 award reviews changed what the government counts as minority health research. https://arxiv.org/abs/2610.03443

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Errors of LLM-Assisted Literature Retrieval in Environmental Science: A Comparison Study of Abstract versus Full-text Based Prompts

Large language models (LLMs) are increasingly used for literature search and synthesis. However, it is unclear whether they retrieve accurate bibliographic information in environmental science. Therefore, we quantitatively compared the errors of widely used LLM platforms in retrieving references related to original articles from five leading environmental science journals (Energy and Environmental Science, Nature Sustainability, Nature Climate Change, Lancet Planetary Health, and Environmental Science and Technology) published in 2024 to 2025. Claude, ChatGPT, Grok, DeepSeek, Perplexity, and Gemini were used as the LLM platforms. LLMs retrieved 10 references for each of the 50 randomly selected original article using either the article's abstract or its full-text as prompt. The retrieved references were subject to a multimetric score ratio combining validity of bibliographic data, Google Scholar link, digital object identifier, Scopus Electronic Identifier and relevance score (cited by or being the index paper), and the proportion of complete fabrication that failed all metrics. Abstract-only prompt yielded significantly higher accuracy than full-text one. This advantage was confirmed in multilevel mixed-effect multivariable regression after adjusting for journal, platform, and output order. Source journal and the position of a reference within the output list were also independently associated with retrieval accuracy, with lower-listed references associated with lower accuracy. These findings suggest that LLM assisted literature retrieval in environmental science remains moderately accurate and overall inconsistent, varying significantly by platform, journal, prompt type, and output position. Abstract-based prompting, as task-aligned information compression, may outperform full-text one in literature retrieval. Caution should be used when generalizing our findings.

cs.DL↗

MTRACE: Multilingual Retrieval-Augmented Generation for Temporally Diverse Text Corpora

Large multilingual knowledge bases expose temporally diverse information, yet retrieval quality remains sensitive to lexical variation and cross-lingual terminology shifts. We develop and evaluate MTRACE (Multilingual Temporal Retrieval-Augmented Generation with evidence grounding), a pipeline designed to test whether query expansion and multi-query fusion mitigate vocabulary mismatch in temporally layered text corpora, on the French and English subsets of MIRACL. Our approach integrates: (i) semantic query expansion (SQE) and multi-query fusion via Reciprocal Rank Fusion (RRF), targeting retrieval stability under query variation; (ii) a generation prompt enforcing strict grounding in retrieved evidence and explicit abstention when evidence is insufficient; and (iii) a modular architecture enabling systematic component evaluation. Ablation studies on Named Entity Recognition (NER) and embedding model selection demonstrate the importance of syntactic coherence in entity extraction and of self-retrieval and efficiency measurements for retriever selection. Our end-to-end evaluation over 50 constructed queries shows faithful answers for well-supported queries, correct abstention on unanswerable questions, and no re-scored similarity gains from multi-query fusion over single-query dense retrieval. By scoping our claims to a clean, text-only baseline, we separate these effects from OCR-noise confounds; direct measurement of diachronic lexical drift is left to future work. Code and configurations are available at \url{https://anonymous.4open.science/r/MIRAGE-8EAA/

cs.DL↗

The machine use of human beings

Research results are prioritised by search engines, large language models and knowledge graphs chiefly through restatement rather than through the originating source. Although the open-access share of annual scholarly output has recently exceeded half, a substantial subscription corpus remains outside what machine readers may legitimately access under prevailing licences. A mechanism is described whereby subscription content is drawn into the open corpus through its citation and restatement in open-access articles---a process termed human mining, by analogy with the text and data mining performed by machines. An optimistic upper bound on what may thereby be inferred by a machine reader confined to the open literature is modelled, and the completeness, latency and licence constraints of that bound are quantified. A majority of subscription articles that have ever been cited is found to be reachable through at least one open citation, and the median interval between a subscription article's publication and its first open citation is shown to have fallen from thirteen years for work of 1990 to one year for work of 2020. The advantage once conferred by direct subscription access to machine readers has thereby been largely eroded, chiefly as a consequence of the growth of open publishing rather than any change in how researchers cite.

cs.DL↗