SearcharxivSearch

arXiv subjects

Daniel Torres-Salinas

Publications and source records attributed to Daniel Torres-Salinas.

At least 19 recordsLinked to original sources

Matching Researchers to Funding Calls: A Reproducible Institution-Level Framework

Grant recommendation systems remain one of the least explored areas within academic recommender systems, and existing proposals are typically tied to specific funding agencies or disciplinary domains. This paper presents an institution-level reproducible framework for matching researchers to funding opportunities by combining bibliometric profiling with semantic matching. Rather than representing each researcher through a single aggregated profile, the framework constructs multiple publication sets defined by bibliometric criteria such as authorship position and time window, each independently compared against funding calls using word embeddings. Within-researcher normalisation and percentile-based ranking transform cosine similarity scores into actionable recommendations. A case study applied to 3,013 researchers from the University of Granada and 291 Horizon Europe topics verify it and shows that the four indicators capture complementary signals.

cs.DL

The 'Big Three' of Scientific Information: A comparative bibliometric review of Web of Science, Scopus, and OpenAlex

The present comparative study examines the three main multidisciplinary bibliographic databases, Web of Science Core Collection, Scopus, and OpenAlex, with the aim of providing up-to-date evidence on coverage, metadata quality, and functional features to help inform strategic decisions in research assessment. The report is structured into two complementary methodological sections. First, it presents a systematic review of recent scholarly literature that investigates record volume, open-access coverage, linguistic diversity, reference coverage, and metadata quality; this is followed by an original bibliometric analysis of the 2015-2024 period that explores longitudinal distribution, document types, thematic profiles, linguistic differences, and overlap between databases. The text concludes with a ten-point executive summary and five recommendations.

cs.DL

Are there stars in Bluesky? A comparative exploratory analysis of altmetric mentions between X and Bluesky

This study examines the shift in the scientific community from X (formerly Twitter) to Bluesky, its impact on scientific communication, and consequently on social metrics (altmetrics). Analyzing 10,174 publications from multidisciplinary and library and information science (LIS) journals in 2024, the results reveal a notable increase in Bluesky activity for multidisciplinary journals in November 2024, likely influenced by political and platform changes, with mentions doubling or quadrupling for journals like Nature and Science. In LIS, the adoption of Bluesky is more limited and shows significant variations across journals, suggesting discipline-specific adoption patterns. However, overall engagement on Bluesky remains significantly lower than on X. While X currently dominates altmetric mentions, the observed growth on Bluesky suggests a potential shift in the future, underscoring its emerging role in academic dissemination and the challenges of adapting scholarly communication metrics across evolving platforms.

cs.DL

The Botization of Science? Large-scale study of the presence and impact of Twitter bots in science dissemination

Twitter bots are a controversial element of the platform, and their negative impact is well known. In the field of scientific communication, they have been perceived in a more positive light, and the accounts that serve as feeds alerting about scientific publications are quite common. However, despite being aware of the presence of bots in the dissemination of science, no large-scale estimations have been made nor has it been evaluated if they can truly interfere with altmetrics. Analyzing a dataset of 3,744,231 papers published between 2017 and 2021 and their associated 51,230,936 Twitter mentions, our goal was to determine the volume of publications mentioned by bots and whether they skew altmetrics indicators. Using the BotometerLite API, we categorized Twitter accounts based on their likelihood of being bots. The results showed that 11,073 accounts (0.23% of total users) exhibited automated behavior, contributing to 4.72% of all mentions. A significant bias was observed in the activity of bots. Their presence was particularly pronounced in disciplines such as Mathematics, Physics, and Space Sciences, with some specialties even exceeding 70% of the tweets. However, these are extreme cases, and the impact of this activity on altmetrics varies by speciality, with minimal influence in Arts & Humanities and Social Sciences. This research emphasizes the importance of distinguishing between specialties and disciplines when using Twitter as an altmetric.

cs.DL

The Many Publics of Science: Using Altmetrics to Identify Common Communication Channels by Scientific field

Altmetrics have led to new quantitative studies of science through social media interactions. However, there are no models of science communication that respond to the multiplicity of non-academic channels. Using the 3653 authors with the highest volume of altmetrics mentions from the main channels (Twitter, News, Facebook, Wikipedia, Blog, Policy documents, and Peer reviews) to their publications (2016-2020), it has been analyzed where the audiences of each discipline are located. The results evidence the generalities and specificities of these new communication models and the differences between areas. These findings are useful for the development of science communication policies and strategies.

cs.DL

Altmetrics can capture research evidence: a study across types of studies in COVID-19 literature

There has been a proliferation of descriptive for COVID-19 papers using altmetrics. The main objective of this study is to analyse whether the altmetric mentions of COVID-19 medical studies are associated with the type of study and its level of evidence. Data were collected from PubMed and Altmetric.com databases. A total of 16,672 study types (e.g., Case reports or Clinical trials) published in the year 2021 and with at least one altmetric mention were retrieved. The altmetric indicators considered were Altmetric Attention Score (AAS), News mentions, Twitter mentions, and Mendeley readers. Once the dataset had been created, the first step was to carry out a descriptive study. Then a normality hypothesis was contrasted by means of the Kolmogorov-Smirnov test, and since it was significant in all cases, the overall comparison of groups was performed using the non-parametric Kruskal-Wallis test. When this test rejected the null hypothesis, pair-by-pair comparisons were performed with the Mann-Whitney U test, and the intensity of the possible association was measured using Cramers V coefficient. The results suggest that the data do not fit a normal distribution. The Mann-Whitney U test revealed coincidences in five groups of study types, the altmetric indicator with most coincidences being news mentions and the study types with the most coincidences were the systematic reviews together with the meta-analyses, which coincided with four altmetric indicators. Likewise, between the study types and the altmetric indicators, a weak but significant association was observed through the chi-square and Cramers V. It is concluded that the positive association between altmetrics and study types in medicine could reflect the level of the pyramid of scientific evidence.

cs.DL

Wikinformetrics: Construction and description of an open Wikipedia knowledge graph dataset for informetric purposes

Wikipedia is one of the most visited websites in the world and is also a frequent subject of scientific research. However, the analytical possibilities of Wikipedia information have not yet been analyzed considering at the same time both a large volume of pages and attributes. The main objective of this work is to offer a methodological framework and an open knowledge graph for the informetric large-scale study of Wikipedia. Features of Wikipedia pages are compared with those of scientific publications to highlight the (di)similarities between the two types of documents. Based on this comparison, different analytical possibilities that Wikipedia and its various data sources offer are explored, ultimately offering a set of metrics meant to study Wikipedia from different analytical dimensions. In parallel, a complete dedicated dataset of the English Wikipedia was built (and shared) following a relational model. Finally, a descriptive case study is carried out on the English Wikipedia dataset to illustrate the analytical potential of the knowledge graph and its metrics.

cs.DL

The growth of COVID-19 scientific literature: A forecast analysis of different daily time series in specific settings

We present a forecasting analysis on the growth of scientific literature related to COVID-19 expected for 2021. Considering the paramount scientific and financial efforts made by the research community to find solutions to end the COVID-19 pandemic, an unprecedented volume of scientific outputs is being produced. This questions the capacity of scientists, politicians and citizens to maintain infrastructure, digest content and take scientifically informed decisions. A crucial aspect is to make predictions to prepare for such a large corpus of scientific literature. Here we base our predictions on the ARIMA model and use two different data sources: the Dimensions and World Health Organization COVID-19 databases. These two sources have the particularity of including in the metadata information the date in which papers were indexed. We present global predictions, plus predictions in three specific settings: type of access (Open Access), NLM source (PubMed and PMC), and domain-specific repository (SSRN and MedRxiv). We conclude by discussing our findings.

cs.DL

The role of scientific output in public debates in times of crisis: A case study of the reopening of schools during the COVID-19 pandemic

Situations in which no scientific consensus has been reached due to either insufficient, inconclusive or contradicting findings place strain on governments and public organizations which are forced to take action under circumstances of uncertainty. In this chapter, we focus on the case of COVID-19, its effects on children and the public debate around the reopening of schools. The aim is to better understand the relationship between policy interventions in the face of an uncertain and rapidly changing knowledge landscape and the subsequent use of scientific information in public debates related to the policy interventions. Our approach is to combine scientific information from journal articles and preprints with their appearance in the popular media, including social media. First, we provide a picture of the different scientific areas and approaches, by which the effects of COVID-19 on children are being studied. Second, we identify news media and social media attention around the COVID-19 scientific output related to children and schools. We focus on policies and media responses in three countries: Spain, South Africa and the Netherlands. These countries have followed very different policy actions with regard to the reopening of schools and represent very different policy approaches to the same problem. We analyse the activity in (social) media around the debate between COVID-19, children and school closures by focusing on the use of references to scientific information in the debate. Finally, we analyse the dominant topics that emerge in the news outlets and the online debates. We draw attention to illustrative cases of miscommunication related to scientific output and conclude the chapter by discussing how information from scientific publication, the media and policy actions shape the public discussion in the context of a global health pandemic.

cs.DL

Exploring WorldCat Identities as an altmetric information source: A library catalog analysis experiment in the field of Scientometrics

Assessing the impact of scholarly books is a difficult research evaluation problem. Library Catalog Analysis facilitates the quantitative study, at different levels, of the impact and diffusion of academic books based on data about their availability in libraries. The WorldCat global catalog collates data on library holdings, offering a range of tools including the novel WorldCat Identities. This is based on author profiles and provides indicators relating to the availability of their books in library catalogs. Here, we investigate this new tool to identify its strengths and weaknesses based on a sample of Bibliometrics and Scientometrics authors. We review the problems that this entails and compare Library Catalog Analysis indicators with Google Scholar and Web of Science citations. The results show that WorldCat Identities can be a useful tool for book impact assessment but the value of its data is undermined by the provision of massive collections of ebooks to academic libraries.

cs.DL

Daily growth rate of scientific production on Covid-19. Analysis in databases and open access repositories

The scientific community is facing one of its greatest challenges in solving a global health problem: COVID-19 pandemic. This situation has generated an unprecedented volume of publications. What is the volume, in terms of publications, of research on COVID-19? The general objective of this research work is to obtain a global vision of the daily growth of scientific production on COVID-19 in different databases (Dimensions, Web of Science Core Collection, Scopus-Elsevier, Pubmed and eight repositories). In relation to the results obtained, Dimensions indexes a total of 9435 publications (69% with peer review and 2677 preprints) well above Scopus (1568) and WoS (718). This is a classic biliometric phenomenon of exponential growth (R2 = 0.92). The global growth rate is 500 publications and the production doubles every 15 days. In the case of Pubmed the weekly growth is around 1000 publications. Of the eight repositories analysed, Pubmed Central, Medrxiv and SSRN are the leaders. Despite their enormous contribution, the journals continue to be the core of scientific communication. Finally, it has been established that three out of every four publications on the COVID-19 are available in open access. The information explosion demands a serious and coordinated response from information professionals, which places us at the centre of the information pandemic.

cs.DL

An alternative analysis on the scientific output of Spanish Sociology What can altmetrics tell us?

In recent years, new indicators known as altmetrics have been introduced to measure the impact of scientific activity. These indicators are obtained through the mentions realised from different social media, existing several aggregators of these data that collect several of them in the same database, being Altmetric.com the most popular. However, in spite of the popularization of these metrics, several limitations in their use have been manifested. For this reason, rhe objective of this work is twofold: (1) to show the possibilities of altimetric techniques applied to the Spanish social sciences in general and sociology in particular; (2) to critically analyse the results to observe the limitations of these indicators; (3) to check whether they can really add useful information that can be used to describe a scientific field and (4) to see the reasons why altmetrics cannot be applied in these fields.

cs.DL

Science through Wikipedia: A novel representation of open knowledge through co-citation networks

This study provides an overview of science from the Wikipedia perspective. A methodology has been established for the analysis of how Wikipedia editors regard science through their references to scientific papers. The method of co-citation has been adapted to this context in order to generate Pathfinder networks (PFNET) that highlight the most relevant scientific journals and categories, and their interactions in order to find out how scientific literature is consumed through this open encyclopaedia. In addition to this, their obsolescence has been studied through Price index. A total of 1 433 457 references available at Altmetric.com have been initially taken into account. After pre-processing and linking them to the data from Elsevier's CiteScore Metrics the sample was reduced to 847 512 references made by 193 802 Wikipedia articles to 598 746 scientific articles belonging to 14 149 journals indexed in Scopus. As highlighted results we found a significative presence of "Medicine" and "Biochemistry, Genetics and Molecular Biology" papers and that the most important journals are multidisciplinary in nature, suggesting also that high-impact factor journals were more likely to be cited. Furthermore, only 13.44% of Wikipedia citations are to Open Access journals.

cs.DL

Library Catalog Analysis and Library Holdings Counts: origins, methodological issues and application to the field of Informetrics

In 2009, Torres-Salinas & Moed proposed the use of library catalogs to analyze the impact and dissemination of academic books in different ways. Library Catalog Analysis (LCA) can be defined as the application of bibliometric techniques to a set of online library catalogs in order to describe quantitatively a scientific-scholarly field on the basis of published book titles. The aim of the present chapter is to conduct an in-depth analysis of major scientific contributions since the birth of LCA in order to determine the state of the art of this research topic. Hence, our specific objectives are: 1) to discuss the original purposes of library holdings 2) to present correlations between library holdings and altmetrics indicators and interpret their feasible meanings 3) to analyze the principal sources of information 4) to use WorldCat Identities to identify the principal authors and works in the field of Informetrics.

cs.DL

Mapping the backbone of the Humanities through the eyes of Wikipedia

The present study aims to establish a valid method by which to apply the theory of co-citations to Wikipedia article references and, subsequently, to map these relationships between scientific papers. This theory, originally applied to scientific literature, will be transferred to the digital environment of collective knowledge generation. To this end, a dataset containing Wikipedia references collected from Altmetric and Scopus' Journal Metrics journals has been used. The articles have been categorized according to the disciplines and specialties established in the All Science Journal Classification (ASJC). They have also been grouped by journal of publication. A set of articles in the Humanities, comprising 25 555 Wikipedia articles with 41 655 references to 32 245 resources, has been selected. Finally, a descriptive statistical study has been conducted and co-citations have been mapped using networks and indicators of degree and betweenness centrality.

cs.DL

Mining university rankings: Publication output and citation impact as their basis

World University rankings have become well-established tools that students, university managers and policy makers read and use. Each ranking claims to have a unique methodology capable of measuring the 'quality' of universities. The purpose of this paper is to analyze to which extent these different rankings measure the same phenomenon and what it is that they are measuring. For this, we selected a total of seven world-university rankings and performed a principal component analysis. After ensuring that despite their methodological differences, they all come together to a single component, we hypothesized that bibliometric indicators could explain what is being measured. Our analyses show that ranking scores from whichever of the seven league tables under study can be explained by the number of publications and citations received by the institution. We conclude by discussing policy implications and opportunities on how a nuanced and responsible use of rankings can help decision making at the institutional level

cs.DL

Mapping social media attention in Microbiology: Identifying main topics and actors

This paper aims to map and identify topics of interest within the field of Microbiology and identify the main sources driving such attention. We combine data from Web of Science and Altmetric.com, a platform which retrieves mentions to scientific literature from social media and other non-academic communication outlets. We focus on the dissemination of microbial publications in Twitter, news media and policy briefs. A two-mode network of social accounts shows distinctive areas of activity. We identify a cluster of papers mentioned solely by regional news media. A central area of the network is formed by papers discussed by the three outlets. A large portion of the network is driven by Twitter activity. When analyzing top actors contributing to such network, we observe that more than half of the Twitter accounts are bots, mentioning 32% of the documents in our dataset. Within news media outlets, there is a predominance of popular science outlets. With regard to policy briefs, both international and national bodies are represented. Finally, our topic analysis shows that the thematic focus of papers mentioned varies by outlet. While news media cover the wider range of topics, policy briefs are focused on translational medicine, and bacterial outbreaks.

cs.DL

The insoluble problems of books: What does Altmetric.com have to offer?

The purpose of this paper is to analyze the capabilities, functionalities and appropriateness of Altmetric.com as a data source for the bibliometric analysis of books in comparison to PlumX. We perform an exploratory analysis on the metrics the Altmetric Explorer for Institutions platform offers for books. We use two distinct datasets of books: the Book Collection included in Altmetric.com and the Clarivate's Master Book List, to analyze Altmetric.com's capabilities to download and merge data with external databases. Finally, we compare our findings with those obtained in a previous study performed in PlumX. Altmetric.com combines and orderly tracks a set of data sources combined by DOI identifiers to retrieve metadata from books, being Google Books its main provider. It also retrieves information from commercial publishers and from some Open Access initiatives, including those led by university libraries such as Harvard Library. We find issues with linkages between records and mentions or ISBN discrepancies. Furthermore, we find that automatic bots affect greatly Wikipedia mentions to books. Our comparison with PlumX suggests that none of these tools provide a complete picture of the social attention generated by books and are rather complementary than comparable tools.

cs.DL