Searcharxiv⌕ Search

arXiv subjects

Enrique Orduna-Malea

Publications and source records attributed to Enrique Orduna-Malea.

At least 19 recordsLinked to original sources

The participation of public in knowledge production: a citizen science projects overview

Citizen Science (CS) is related to public engagement in scientific research. The tasks in which the citizens can be involved are diverse and can range from data collection and tagging images to participation in the planning and research design. However, little is known about the involvement degree of the citizens to CS projects, and the contribution of those projects to the advancement of knowledge (e.g. scientific outcomes). This study aims to gain a better understanding by analysing the SciStarter database. A total of 2,346 CS projects were identified, mainly from Ecology and Environmental Sciences. Of these projects, 91% show low participation of the citizens (Level 1 "citizens as sensors" and 2 "citizens as interpreters", from Haklay's scale). In terms of scientific output, 918 papers indexed in the Web of Science (WoS) were identified. The most prolific projects were found to have lower levels of citizen involvement, specifically at Levels 1 and 2.

cs.DL↗

A scientometric-inspired framework to analyze EurekAlert! press releases

Press releases about scholarly news are brief statements provided in advance to the press, including a description of the most relevant findings of one or more accepted scientific publications, usually under the condition that journalists will adhere to an embargo until the publication date. The existence of centralized platforms such as EurekAlert! allows press releases to be disseminated online as independent news articles. Press releases can include additional material (e.g., interviews, commentaries, explanatory tables, figures, media, recommended readings), which turn them into online objects with analytical value of their own. The objective of this work is to illustrate how press releases can be quantitatively analyzed applying similar tools and approaches as those applied in scientometric research (SCI). To achieve this goal, a scientometric inspired analytical framework is proposed based on the formulation of spaces of interaction of objects, actors, and impacts. As such, the framework proposed considers press releases as science communication (SCO) objects, produced by different SCO actors (e.g., journalists), and the subject of receiving impact (e.g., tweets, links). To carry out this analysis, all press releases published by EurekAlert! from 1996 until 2021 (455,703 press releases), all tweets including at least one URL referring to a EurekAlert! press release (1,364,563 tweets), and all webpages with at least one URL referring to a EurekAlert! press release (54,089,233 webpages) have been studied. We argue that the large volume of press releases published and their online dissemination make these objects relevant in the measurement of SCO-SCI interactions.

cs.DL↗

Measuring web connectivity between research organizations through ROR identifiers

Digital information needs to be accessed and used in a manageable and sustainable manner to facilitate the advancement of science and science management. Many types of Persistent Identifiers (PIDs) are already in use and well-established in support of the scholarly communication industry, mainly digital objects (e.g., DOIs) and person identifiers (e.g., ORCID). PIDs improve the interoperability of digital entities, make them reusable, and, at the same time, foster FAIR principles. The main objective of this exploratory work is to measure the degree and type of use of ROR identifiers by the online scientific and academic ecosystem through link-based indicators. The analysis yielded 149,851 links to ror.org webpages: 147,154 links to ROR-based URLs and 2,698 links to other informative webpages under the ror.org website. The results obtained evidence that the percentage of ROR identifiers linked is limited (51.6% of ROR identifiers have been linked at least once). These links come from a limited number of referring domains (242 unique domain names) and mainly from bibliographic records (51.4% of links) and organization cards (36% of links). While the distribution of ROR identifiers is biased towards Anglo-Saxon countries (mainly United States) and types (companies), the educational research organizations are the institutions most linked through their corresponding ROR-based URLs. The connectivity between DOIs, ORCIDs and RORs can be the spearhead to carry out new webometric and bibliometric studies, of interest to characterize the presence, impact, and interconnection of the global academic Web.

cs.DL↗

Dot-Science Top Level Domain: academic websites or dumpsites?

Dot-science was launched in 2015 as a new academic top-level domain (TLD) aimed to provide 'a dedicated, easily accessible location for global Internet users with an interest in science'. The main objective of this work is to find out the general scholarly usage of this top-level domain. In particular, the following three questions are pursued: usage (number of web domains registered with the dot-science), purpose (main function and category of websites linked to these web domains), and impact (websites' visibility and authority). To do this, 13,900 domain names were gathered through ICANN's Domain Name Registration Data Lookup database. Each web domain was subsequently categorized, and data on web impact were obtained from Majestic's API. Based on the results obtained, it is concluded that the dot-science top-level domain is scarcely adopted by the academic community, and mainly used by registrar companies for reselling purposes (35.5% of all web domains were parked). Websites receiving the highest number of backlinks were generally related to non-academic websites applying intensive link building practices and offering leisure or even fraudulent contents. Majestic's Trust Flow metric has been proved an effective method to filter reputable academic websites. As regards primary academic-related dot-science web domain categories, 1,175 (8.5% of all web domains registered) were found, mainly personal academic websites (342 web domains), blogs (261) and research groups (133). All dubious content reveals bad practices on the Web, where the tag 'science' is fundamentally used as a mechanism to deceive search engine algorithms.

cs.DL↗

Google Scholar, Microsoft Academic, Scopus, Dimensions, Web of Science, and OpenCitations' COCI: a multidisciplinary comparison of coverage via citations

New sources of citation data have recently become available, such as Microsoft Academic, Dimensions, and the OpenCitations Index of CrossRef open DOI-to-DOI citations (COCI). Although these have been compared to the Web of Science (WoS), Scopus, or Google Scholar, there is no systematic evidence of their differences across subject categories. In response, this paper investigates 3,073,351 citations found by these six data sources to 2,515 English-language highly-cited documents published in 2006 from 252 subject categories, expanding and updating the largest previous study. Google Scholar found 88% of all citations, many of which were not found by the other sources, and nearly all citations found by the remaining sources (89%-94%). A similar pattern held within most subject categories. Microsoft Academic is the second largest overall (60% of all citations), including 82% of Scopus citations and 86% of Web of Science citations. In most categories, Microsoft Academic found more citations than Scopus and WoS (182 and 223 subject categories, respectively), but had coverage gaps in some areas, such as Physics and some Humanities categories. After Scopus, Dimensions is fourth largest (54% of all citations), including 84% of Scopus citations and 88% of WoS citations. It found more citations than Scopus in 36 categories, more than WoS in 185, and displays some coverage gaps, especially in the Humanities. Following WoS, COCI is the smallest, with 28% of all citations. Google Scholar is still the most comprehensive source. In many subject categories Microsoft Academic and Dimensions are good alternatives to Scopus and WoS in terms of coverage.

cs.DL↗

Universities through the Eyes of Bibliographic Databases: A Retroactive Growth Comparison of Google Scholar, Scopus and Web of Science

The purpose of this study is to ascertain the suitability of GS's url-based method as a valid approximation of universities' academic output measures, taking into account three aspects (retroactive growth, correlation, and coverage). To do this, a set of 100 Turkish universities were selected as a case study. The productivity in Web of Science (WoS), Scopus and GS (2000 to 2013) were captured in two different measurement iterations (2014 and 2018). In addition, a total of 18,174 documents published by a subset of 14 research-focused universities were retrieved from WoS, verifying their presence in GS within the official university web domain. Findings suggest that the retroactive growth in GS is unpredictable and dependent on each university, making this parameter hard to evaluate at the institutional level. Otherwise, the correlation of productivity between GS (url-based method) and WoS and Scopus (selected sources) is moderately positive, even though it varies depending on the university, the year of publication, and the year of measurement. Finally, only 16% out of 18,174 articles analyzed were indexed in the official university website, although up to 84% were indexed in other GS sources. This work proves that the url-based method to calculate institutional productivity in GS is not a good proxy for the total number of publications indexed in WoS and Scopus, at least in the national context analyzed. However, the main reason is not directly related to the operation of GS, but with a lack of universities' commitment to open access.

cs.DL↗

Crossing the Academic Ocean? Judit Bar-Ilan's Oeuvre on Search Engines studies

The main objective of this work is to analyse the contributions of Judit Bar-Ilan to the search engines studies. To do this, two complementary approaches have been carried out. First, a systematic literature review of 47 publications authored and co-authored by Judit and devoted to this topic. Second, an interdisciplinarity analysis based on the cited references (publications cited by Judit) and citing documents (publications that cite Judit's work) through Scopus. The systematic literature review unravels an immense amount of search engines studied (43) and indicators measured (especially technical precision, overlap and fluctuation over time). In addition to this, an evolution over the years is detected from descriptive statistical studies towards empirical user studies, with a mixture of quantitative and qualitative methods. Otherwise, the interdisciplinary analysis evidences that a significant portion of Judit's oeuvre was intellectually founded on the computer sciences, achieving a significant, but not exclusively, impact on library and information sciences.

cs.DL↗

Google Scholar, Web of Science, and Scopus: a systematic comparison of citations in 252 subject categories

Despite citation counts from Google Scholar (GS), Web of Science (WoS), and Scopus being widely consulted by researchers and sometimes used in research evaluations, there is no recent or systematic evidence about the differences between them. In response, this paper investigates 2,448,055 citations to 2,299 English-language highly-cited documents from 252 GS subject categories published in 2006, comparing GS, the WoS Core Collection, and Scopus. GS consistently found the largest percentage of citations across all areas (93%-96%), far ahead of Scopus (35%-77%) and WoS (27%-73%). GS found nearly all the WoS (95%) and Scopus (92%) citations. Most citations found only by GS were from non-journal sources (48%-65%), including theses, books, conference papers, and unpublished materials. Many were non-English (19%-38%), and they tended to be much less cited than citing sources that were also in Scopus or WoS. Despite the many unique GS citing sources, Spearman correlations between citation counts in GS and WoS or Scopus are high (0.78-0.99). They are lower in the Humanities, and lower between GS and WoS than between GS and Scopus. The results suggest that in all areas GS citation data is essentially a superset of WoS and Scopus, with substantial extra coverage.

cs.DL↗

Do the technical universities exhibit distinct behaviour in global university rankings? A Times Higher Education (THE) case study

Technical Universities (TUs) exhibit a distinct ranking performance in comparison with other universities. In this paper we identify 137 TUs included in the THE Ranking (2017 edition) and analyse their scores statistically. The results highlight the existence of clusters of TUs showing a general high performance in the Industry Income category and, in many cases, a low performance on Research and Teaching. Finally, the global score weights were simulated, creating several scenarios that confirmed that the majority of TUs (except those with a world-class status) would increase their final scores if industrial income was accounted for at the levels parametrised.

cs.DL↗

Coverage of highly-cited documents in Google Scholar, Web of Science, and Scopus: a multidisciplinary comparison

This study explores the extent to which bibliometric indicators based on counts of highly-cited documents could be affected by the choice of data source. The initial hypothesis is that databases that rely on journal selection criteria for their document coverage may not necessarily provide an accurate representation of highly-cited documents across all subject areas, while inclusive databases, which give each document the chance to stand on its own merits, might be better suited to identify highly-cited documents. To test this hypothesis, an analysis of 2,515 highly-cited documents published in 2006 that Google Scholar displays in its Classic Papers product is carried out at the level of broad subject categories, checking whether these documents are also covered in Web of Science and Scopus, and whether the citation counts offered by the different sources are similar. The results show that a large fraction of highly-cited documents in the Social Sciences and Humanities (8.6%-28.2%) are invisible to Web of Science and Scopus. In the Natural, Life, and Health Sciences the proportion of missing highly-cited documents in Web of Science and Scopus is much lower. Furthermore, in all areas, Spearman correlation coefficients of citation counts in Google Scholar, as compared to Web of Science and Scopus citation counts, are remarkably strong (.83-.99). The main conclusion is that the data about highly-cited documents available in the inclusive database Google Scholar does indeed reveal significant coverage deficiencies in Web of Science and Scopus in several areas of research. Therefore, using these selective databases to compute bibliometric indicators based on counts of highly-cited documents might produce biased assessments in poorly covered areas.

cs.DL↗

Google Scholar as a data source for research assessment

The launch of Google Scholar (GS) marked the beginning of a revolution in the scientific information market. This search engine, unlike traditional databases, automatically indexes information from the academic web. Its ease of use, together with its wide coverage and fast indexing speed, have made it the first tool most scientists currently turn to when they need to carry out a literature search. Additionally, the fact that its search results were accompanied from the beginning by citation counts, as well as the later development of secondary products which leverage this citation data (such as Google Scholar Metrics and Google Scholar Citations), made many scientists wonder about its potential as a source of data for bibliometric analyses. The goal of this chapter is to lay the foundations for the use of GS as a supplementary source (and in some disciplines, arguably the best alternative) for scientific evaluation. First, we present a general overview of how GS works. Second, we present empirical evidences about its main characteristics (size, coverage, and growth rate). Third, we carry out a systematic analysis of the main limitations this search engine presents as a tool for the evaluation of scientific performance. Lastly, we discuss the main differences between GS and other more traditional bibliographic databases in light of the correlations found between their citation data. We conclude that Google Scholar presents a broader view of the academic world because it has brought to light a great amount of sources that were not previously visible.

cs.DL↗

Google Scholar: the 'big data' bibliographic tool

The launch of Google Scholar back in 2004 meant a revolution not only in the scientific information search market but also in research evaluation processes. Its dynamism, unparalleled coverage, and uncontrolled indexing make of Google Scholar an unusual product, especially when compared to traditional bibliographic databases. Conceived primarily as a discovery tool for academic information, it presents a number of limitations as a bibliometric tool. The main objective of this chapter is to show how Google Scholar operates and how its core database may be used for bibliometric purposes. To do this, the general features of the search engine (in terms of document typologies, disciplines, and coverage) are analysed. Lastly, several bibliometric tools based on Google Scholar data, both official (Google Scholar Metrics, Google Scholar Citations), and some developed by third parties (H Index Scholar, Publishers Scholar Metrics, Proceedings Scholar Metrics, Journal Scholar Metrics, Scholar Mirrors), as well as software to collect and process data from this source (Publish or Perish, Scholarometer) are introduced, aiming to illustrate the potential bibliometric uses of this source.

cs.DL↗

A novel method for depicting academic disciplines through Google Scholar Citations: The case of Bibliometrics

This article describes a procedure to generate a snapshot of the structure of a specific scientific community and their outputs based on the information available in Google Scholar Citations (GSC). We call this method MADAP (Multifaceted Analysis of Disciplines through Academic Profiles). The international community of researchers working in Bibliometrics, Scientometrics, Informetrics, Webometrics, and Altmetrics was selected as a case study. The records of the top 1,000 most cited documents by these authors according to GSC were manually processed to fill any missing information and deduplicate fields like the journal titles and book publishers. The results suggest that it is feasible to use GSC and the MADAP method to produce an accurate depiction of the community of researchers working in Bibliometrics (both specialists and occasional researchers) and their publication habits (main publication venues such as journals and book publishers). Additionally, the wide document coverage of Google Scholar (specially books and book chapters) enables more comprehensive analyses of the documents published in a specific discipline than were previously possible with other citation indexes, finally shedding light on what until now had been a blind spot in most citation analyses.

cs.DL↗

Can we use Google Scholar to identify highly-cited documents?

The main objective of this paper is to empirically test whether the identification of highly-cited documents through Google Scholar is feasible and reliable. To this end, we carried out a longitudinal analysis (1950 to 2013), running a generic query (filtered only by year of publication) to minimise the effects of academic search engine optimisation. This gave us a final sample of 64,000 documents (1,000 per year). The strong correlation between a document's citations and its position in the search results (r= -0.67) led us to conclude that Google Scholar is able to identify highly-cited papers effectively. This, combined with Google Scholar's unique coverage (no restrictions on document type and source), makes the academic search engine an invaluable tool for bibliometric research relating to the identification of the most influential scientific documents. We find evidence, however, that Google Scholar ranks those documents whose language (or geographical web domain) matches with the user's interface language higher than could be expected based on citations. Nonetheless, this language effect and other factors related to the Google Scholar's operation, i.e. the proper identification of versions and the date of publication, only have an incidental impact. They do not compromise the ability of Google Scholar to identify the highly-cited papers.

cs.DL↗

Author-level metrics in the new academic profile platforms: The online behaviour of the Bibliometrics community

The new web-based academic communication platforms do not only enable researchers to better advertise their academic outputs, making them more visible than ever before, but they also provide a wide supply of metrics to help authors better understand the impact their work is making. This study has three objectives: a) to analyse the uptake of some of the most popular platforms (Google Scholar Citations, ResearcherID, ResearchGate, Mendeley and Twitter) by a specific scientific community (bibliometrics, scientometrics, informetrics, webometrics, and altmetrics); b) to compare the metrics available from each platform; and c) to determine the meaning of all these new metrics. To do this, the data available in these platforms about a sample of 811 authors (researchers in bibliometrics for whom a public profile Google Scholar Citations was found) were extracted. A total of 31 metrics were analysed. The results show that a high number of the analysed researchers only had a profile in Google Scholar Citations (159), or only in Google Scholar Citations and ResearchGate (142). Lastly, we find two kinds of metrics of online impact. First, metrics related to connectivity (followers), and second, all metrics associated to academic impact. This second group can further be divided into usage metrics (reads, views), and citation metrics. The results suggest that Google Scholar Citations is the source that provides more comprehensive citation-related data, whereas Twitter stands out in connectivity-related metrics.

cs.DL↗

Dimensions: re-discovering the ecosystem of scientific information

The overarching aim of this work is to provide a detailed description of the free version of Dimensions (new bibliographic database produced by Digital Science and launched in January 2018). To do this, the work is divided into two differentiated blocks. First, its characteristics, operation and features are described, focusing on its main strengths and weaknesses. Secondly, an analysis of its coverage is carried out (comparing it Scopus and Google Scholar) in order to determine whether the bibliometric indicators offered by Dimensions have an order of magnitude significant enough to be used. To this end, an analysis is carried out at three levels: journals (sample of 20 publications in 'Library & Information Science'), documents (276 articles published by the Journal of informetrics between 2013 and 2015) and authors (28 people awarded with the Derek de Solla Price prize). Preliminary results indicate that Dimensions has coverage of the recent literature superior to Scopus although inferior to Google Scholar. With regard to the number of citations received, Dimensions offers slightly lower figures than Scopus. Despite this, the number of citations in Dimensions exhibits a strong correlation with Scopus and somewhat less (although still significant) with Google Scholar. For this reason, it is concluded that Dimensions is an alternative for carrying out citation studies, being able to rival Scopus (greater coverage and free of charge) and with Google Scholar (greater functionalities for the treatment and data export).

cs.DL↗

Do ResearchGate Scores create ghost academic reputations?

The academic social network site ResearchGate (RG) has its own indicator, RG Score, for its members. The high profile nature of the site means that the RG score may be used for recruitment, promotion and other tasks for which researchers are evaluated. In response, this study investigates whether it is reasonable to employ the RG Score as evidence of scholarly reputation. For this, three different author samples were investigated. An outlier sample includes 104 authors with high values. A Nobel sample comprises 73 Nobel winners from Medicine & Physiology, Chemistry, Physics and Economics (from 1975 to 2015). A longitudinal sample includes weekly data on 4 authors with different RG Scores. The results suggest that high RG Scores are built primarily from activity related to asking and answering questions in the site. In particular, it seems impossible to get a high RG Score solely through publications. Within RG it is possible to distinguish between (passive) academics that interact little in the site and active platform users, who can get high RG Scores through engaging with others inside the site (questions, answers, social networks with influential researchers). Thus, RG Scores should not be mistaken for academic reputation indicators.

cs.SI↗

Google Scholar and the gray literature: A reply to Bonato's review

Recently, a review concluded that Google Scholar (GS) is not a suitable source of information "for identifying recent conference papers or other gray literature publications". The goal of this letter is to demonstrate that GS can be an effective tool to search and find gray literature, as long as appropriate search strategies are used. To do this, we took as examples the same two case studies used by the original review, describing first how GS processes original's search strategies, then proposing alternative search strategies, and finally generalizing each case study to compose a general search procedure aimed at finding gray literature in Google Scholar for two wide selected case studies: a) all contributions belonging to a congress (the ASCO Annual Meeting); and b) indexed guidelines as well as gray literature within medical institutions (National Institutes of Health) and governmental agencies (U.S. Department of Health & Human Services). The results confirm that original search strategies were undertrained offering misleading results and erroneous conclusions. Google Scholar lacks many of the advanced search features available in other bibliographic databases (such as Pubmed), however, it is one thing to have a friendly search experience, and quite another to find gray literature. We finally conclude that Google Scholar is a powerful tool for searching gray literature, as long as the users are familiar with all the possibilities it offers as a search engine. Poorly formulated searches will undoubtedly return misleading results.

cs.DL↗