Searcharxiv⌕ Search

arXiv subjects

Maxime Holmberg Sainte-Marie

Publications and source records attributed to Maxime Holmberg Sainte-Marie.

3 recordsLinked to original sources

Any Old Tom, Dick or Harry: The Citation Impact of First Name Genderedness

This paper examines the relationship between the genderedness of authors' first names and citation distributions in US-affiliated scholarly production from 2010 to 2019. Merging a first name genderedness table derived from Wikidata with Web of Science publication data, we develop a relative distributional framework that compares name, article, and citation distributions along a continuous genderedness spectrum. The lexical structure of the corpus proves stable across author roles, while productivity diverges substantially by disciplinary group: physical sciences show a consistent masculine lean, social sciences a feminine one, with amplitude remaining stable across groups despite these directional differences. Citation analyses net of publication volume reveal a pervasive deficit across disciplinary groups and author roles: only unambiguously masculine names accumulate citations in excess of their publication volume, while feminine, neutral, and moderately masculine names all fall short of it. This asymmetry is strongest in the life sciences and weakest in the physical sciences except among middle authors, where the pattern reflects the gender composition of large collaborative teams rather than evaluative bias in citing behavior. Taken together, these results are consistent with the hypothesis that first name genderedness shapes citation recognition through implicit bias operating in low-deliberation evaluative contexts.

cs.DL↗

Sorting the Babble in Babel: Assessing the Performance of Language Identification Algorithms on the OpenAlex Database

This project aims to optimize the linguistic indexing of the OpenAlex database by comparing the performance of various Python-based language identification procedures on different metadata corpora extracted from a manually-annotated article sample \footnote{OpenAlex used the results presented in this article to inform the language metadata overhaul carried out as part of its recent Walden system launch. The precision and recall performance of each algorithm, corpus, and language is first analyzed, followed by an assessment of processing speeds recorded for each algorithm and corpus type. These different performance measures are then simulated at the database level using probabilistic confusion matrices for each algorithm, corpus, and language, as well as a probabilistic modeling of relative article language frequencies for the whole OpenAlex database. Results show that procedure performance strongly depends on the importance given to each of the measures implemented: for contexts where precision is preferred, using the LangID algorithm on the greedy corpus gives the best results; however, for all cases where recall is considered at least slightly more important than precision or as soon as processing times are given any kind of consideration, the procedure that consists in the application of the FastText algorithm on the Titles corpus outperforms all other alternatives. Given the lack of truly multilingual large-scale bibliographic databases, it is hoped that these results help confirm and foster the unparalleled potential of the OpenAlex database for cross-linguistic and comprehensive measurement and evaluation.

cs.CL↗

Evaluating the Linguistic Coverage of OpenAlex: An Assessment of Metadata Accuracy and Completeness

Clarivate's Web of Science (WoS) and Elsevier's Scopus have been for decades the main sources of bibliometric information. Although highly curated, these closed, proprietary databases are largely biased towards English-language publications, underestimating the use of other languages in research dissemination. Launched in 2022, OpenAlex promised comprehensive, inclusive, and open-source research information. While already in use by scholars and research institutions, the quality of its metadata is currently being assessed. This paper contributes to this literature by assessing the completeness and accuracy of OpenAlex's metadata related to language, through a comparison with WoS, as well as an in-depth manual validation of a sample of 6,836 articles. Results show that OpenAlex exhibits a far more balanced linguistic coverage than WoS. However, language metadata is not always accurate, which leads OpenAlex to overestimate the place of English while underestimating that of other languages. If used critically, OpenAlex can provide comprehensive and representative analyses of languages used for scholarly publishing. However, more work is needed at infrastructural level to ensure the quality of metadata on language.

cs.DL↗