SearcharxivSearch

arXiv subjects

Fidan Badalova

Publications and source records attributed to Fidan Badalova.

3 recordsLinked to original sources

Detecting Hallucinated and Suspicious Citations: What Current Tools Can and Cannot Do

Large language models are increasingly used in academic writing, including for reference generation, raising concerns about hallucinated and unreliable citations. Recent research suggests that this problem is already widespread and is becoming increasingly prevalent in the published literature and at scientific conferences. In this position paper, we review recent studies on hallucinated references and evaluate several currently available tools for detecting problematic references using documents containing hallucinated citations. The tools assessed include CheckIfExist, HalluCiteChecker, Hallucinator, Hallucinated Reference Finder (HalRef), and RefChecker. While these systems can provide useful early warnings in many cases, their performance is limited by reference extraction errors, incomplete metadata, limited database coverage, and inconsistent verification results. We argue that hallucinated and suspicious references have become a real and growing problem for scientific communication, and that more transparent and multi-source detection systems are still needed.

cs.DL

PreprintToPaper dataset: connecting bioRxiv preprints with journal publications

The PreprintToPaper dataset connects bioRxiv preprints with their corresponding journal publications, enabling large-scale analysis of the preprint-to-publication process. It comprises metadata for 145,517 preprints from two periods, 2016-2018 (pre-pandemic) and 2020-2022 (pandemic), retrieved via the bioRxiv and Crossref APIs. We selected the two periods to capture preprint-publication dynamics before and during the COVID-19 pandemic while avoiding transitional years. Each record includes bibliographic information such as titles, abstracts, authors, institutions, submission dates, licenses, and subject categories, alongside enriched publication metadata including journal names, publication dates, author lists, and further information. In addition to the main dataset, a version-history subset provides all available versions of preprints within the two selected periods, enabling analysis of how preprints evolve over time. Preprints are categorized into three groups: Published (formally linked to a journal article), Preprint Only (posted on a preprint server), and Gray Zone (potentially published in a journal but unlinked). To enhance reliability, title and author similarity scores were computed, and a human-annotated subset of 299 records was created to evaluate Gray Zone cases. The dataset supports diverse applications, including studies of scholarly communication, open science policies, bibliometric tool development, and natural language processing research on textual changes between preprints and the corresponding journal articles. The dataset is publicly available in CSV format via Zenodo.

cs.DL

Unveiling tortured phrases in Humanities and Social Sciences

A small amount of unscrupulous people, concerned by their career prospects, resort to paper mill services to publish articles in renowned journals and conference proceedings. These include patchworks of synonymized contents using paraphrasing tools, featuring tortured phrases, increasingly polluting the scientific literature. The Problematic Paper Screener (PPS) has been developed to allow articles (re)assessment on PubPeer. Since most of the known tortured phrases are found in publications in science, technology, engineering, and mathematics (STEM), we extend this work by exploring their presence in the humanities and social sciences (HSS). To do so, we used the PPS to look for tortured abbreviations, generated from the two social science thesauri ELSST and THESOZ. We also used two case studies to find new tortured abbreviations, by screening the Hindawi EDRI journal and the GESIS SSOAR repository. We found a total of 32 multidisciplinary problematic documents, related to Education, Psychology, and Economics. We also generated 121 new fingerprints to be added to the PPS. These articles and future screening have to be investigated by social scientists, as most of it is currently done by STEM domain experts.

cs.DL