SearcharxivSearch

arXiv subjects

Zhichao Fang

Publications and source records attributed to Zhichao Fang.

At least 19 recordsLinked to original sources

Should children follow their parents' research paths? Intergenerational research continuity and divergence in academic families

How academic advantages are transmitted within families is usually studied as occupational inheritance, but it is not clear whether scholarly research orientations persist across generations and if it is an advantage when it does. To address this, we link Wikidata kinship records with OpenAlex bibliometric profiles to study 3,229 documented parent-child scholar pairs and 488,659 publications. Field-level research similarity was evident but not universal: whilst the median similarity was 0.546, 25.3% of parent-child pairs had no Field overlap (i.e., similarity 0). These pairs were substantially more similar than publication-period-matched comparison pairs (median 0.098). Direct academic interaction was uncommon: 10.4% of parent-child pairs had co-authored, 9.8% of children had cited their parents, and 6.9% of parents had cited their children. Nevertheless, each 0.1 increase in Field similarity was associated with 38-39% higher adjusted odds of co-authorship and cross-citation. There was also intergenerational continuity in academic achievement and recognition. Parents' publication volume and field-normalized citation impact were positively associated with those of their children. Children of national academy members had approximately twice the odds of becoming national academy members themselves (Odds Ratio = 2.04), while children of prizewinning parents had 46% higher odds of winning prizes (Odds Ratio = 1.46). However, children of national academy members showed lower research similarity to their parents. Greater research differentiation was associated with higher field-normalized citation impact among children, but not with publication output or higher odds of academic recognition. Academic families therefore appear to transmit resources and advantages with the sole exception that diverging from parental fields seems to confer a citation advantage.

cs.DL

Science discussions of retracted articles on Bluesky: public scrutiny or misinformation spreading?

Post-publication peer review (PPPR) has emerged as an important supplement to traditional peer review, with social media playing a growing role in publicising potential problems in published research. However, it remains unclear whether social media discussions of retracted articles primarily reflect good practices, such as exposing flaws and acknowledging retraction status, or bad practices, such as overlooking retractions and continuing to disseminate scientific misinformation. In this study, we collected Bluesky posts referencing scholarly articles from Altmetric and retrieved metadata for the referenced articles using OpenAlex. The final dataset included 284 retracted articles with 79 pre-retraction posts and 857 post-retraction posts, 59 retraction notices with 186 posts, and 609,461 non-retracted articles with 1,344,756 posts. We manually coded Bluesky posts discussing retracted articles to identify instances of good and bad practice. The results show that posts demonstrating good practice (89.9%) substantially outnumbered those demonstrating bad practice (10.1%). Posts reflecting good practice also had more user engagement. In the pre-retraction phase, good practice posts constituted a slight minority (43.0%), whereas in the post-retraction phase they were dominant (94.2%). Most negative posts in the pre-retraction phase (90.0%) had good practice while only 17.3% positive posts in the post-retraction phase showed bad practice. Thus, sentiment analysis can be helpful to filter posts that could flag potential flaws before retraction, but it may struggle to accurately identify the spread of misinformation after retraction. More broadly, this study highlights the potential of Bluesky to support responsible scientific communication, public scrutiny, and research integrity.

cs.DL

Organisational accounts engaged in scholarly communication on Twitter: Patterns of presence, activity and engagement

Organisational accounts are an integral part of the Twitter (now X) ecosystem. This study identified 9,842 research- and policy-related organisational accounts that had tweeted about scholarly publications by linking three global organisational databases (GRID, ROR, and Overton) with two altmetric databases containing Twitter data (Altmetric and the former Crossref Event Data). The resulting openly available dataset was used to examine organisational activity in scholarly communication across three dimensions: social media capital, tweeting activity, and engagement level. The results show that, compared to all Twitter users engaged in scholarly communication, organisational accounts hold a notable advantage in terms of follower bases and the proportion of scholarly tweets. Their scholarly tweets achieve high visibility through likes and retweets but perform weakly in generating more conversational forms of engagement, such as quotes and replies. Distinct patterns emerge across organisational categories: research facilities, in particular, demonstrate the strongest focus on scholarly tweeting, whereas government accounts are comparatively more successful in eliciting engagement across all metrics, including the more interactive ones. This study contributes both an open dataset of organisational accounts and a methodological framework for their identification, while also highlighting the important roles that organisations play in shaping scholarly discourse on social media.

cs.DL

Evidence for studying interactions between science and policy: An exploration of scholarly and policy references in Overton-indexed policy documents

Overton, a global policy index, provides new opportunities to study the interactions between science and policy. This study aims to characterize the presence of scholarly and policy references in Overton-indexed policy documents and examine their distribution across key bibliographic dimensions, thereby assessing Overton's potential as a data source for policy metrics. We analyze a dataset of approximately 17.5 million policy documents from Overton, incorporating metadata such as publication year, policy source, country, language, subject area, and policy topic. Descriptive statistics are employed to assess the presence and distribution of reference data across these dimensions. Overton indexes a substantial volume of policy documents and identifies considerable reference data within them: 7.7% of documents contain scholarly references and 10.6% contain policy references. However, the presence of references varies significantly across publications years, source types, countries, languages, subject areas, and policy topics, indicating coverage biases that may affect interpretations of policy impact. The analysis is based on the Overton database as of June 2025. As Overton is regularly updated, the distribution patterns of indexed documents and references may evolve over time. The findings offer insights into the opportunities and constraints of using Overton for investigating evidence-based policymaking and for assessing the policy uptake of research outputs in the context of research evaluation. This is the first large-scale study to systematically examine the distribution of reference data in Overton. It contributes a foundational understanding of this emerging source for policy metrics, highlighting both its potential applications and limitations, and underlining the importance of addressing current coverage imbalances.

cs.DL

Can social media provide early warning of retraction? Evidence from critical tweets identified by human annotation and large language models

Timely detection of problematic research is essential for safeguarding scientific integrity. To explore whether social media commentary can serve as an early indicator of potentially problematic articles, this study analysed 3,815 tweets referencing 604 retracted articles and 3,373 tweets referencing 668 comparable non-retracted articles. Tweets critical of the articles were identified through both human annotation and large language models (LLMs). Human annotation revealed that 8.3% of retracted articles were associated with at least one critical tweet prior to retraction, compared to only 1.5% of non-retracted articles, highlighting the potential of tweets as early warning signals of retraction. However, critical tweets identified by LLMs (GPT-4o mini, Gemini 2.0 Flash-Lite, and Claude 3.5 Haiku) only partially aligned with human annotation, suggesting that fully automated monitoring of post-publication discourse should be applied with caution. A human-AI collaborative approach may offer a more reliable and scalable alternative, with human expertise helping to filter out tweets critical of issues unrelated to the research integrity of the articles. Overall, this study provides insights into how social media signals, combined with generative AI technologies, may support efforts to strengthen research integrity.

cs.DL

How is science discussed on Bluesky?

Amid the migration of academics from X, the social media platform Bluesky has emerged as a potential alternative. To assess its viability and relevance for science communication, this study presents the first large-scale analysis of scholarly article dissemination on Bluesky, exploring its potential as a new source of social media metrics. We collected and analysed over 2.6 million Bluesky posts referencing 532,302 scholarly articles from January 2023 to July 2025, integrating metadata from the OpenAlex database. Temporal trends, disciplinary coverage, language use, textual characteristics, and user engagement were examined. A sharp increase in scholarly activity on Bluesky was observed from November 2024 to January 2025, coinciding with broader academic shifts away from X. As on X, Bluesky posts primarily concern the health, social, and environmental sciences and are predominantly written in English. Nevertheless, Bluesky posts demonstrate substantially higher levels of interaction (likes, reposts, replies, and quotes) and greater textual originality than previously reported for X, suggesting both stronger interactive and more interpretive engagement. These findings highlight Bluesky's emerging role as a credible platform for science communication and a promising source for altmetrics. The platform may facilitate not only early visibility of research outputs but also more meaningful scholarly dialogue in the evolving social media landscape.

cs.DL

Exploring the Technical Knowledge Interaction of Global Digital Humanities: Three-decade Evidence from Bibliometric-based perspectives

Digital Humanities (DH) is an interdisciplinary field that integrates computational methods with humanities scholarship to investigate innovative topics. Each academic discipline follows a unique developmental path shaped by the topics researchers investigate and the methods they employ. With the help of bibliometric analysis, most of previous studies have examined DH across multiple dimensions such as research hotspots, co-author networks, and institutional rankings. However, these studies have often been limited in their ability to provide deep insights into the current state of technological advancements and topic development in DH. As a result, their conclusions tend to remain superficial or lack interpretability in understanding how methods and topics interrelate in the field. To address this gap, this study introduced a new concept of Topic-Method Composition (TMC), which refers to a hybrid knowledge structure generated by the co-occurrence of specific research topics and the corresponding method. Especially by analyzing the interaction between TMCs, we can see more clearly the intersection and integration of digital technology and humanistic subjects in DH. Moreover, this study developed a TMC-based workflow combining bibliometric analysis, topic modeling, and network analysis to analyze the development characteristics and patterns of research disciplines. By applying this workflow to large-scale bibliometric data, it enables a detailed view of the knowledge structures, providing a tool adaptable to other fields.

cs.DL

Do male leading authors retract more articles than female leading authors?

Scientific retractions reflect issues within the scientific record, arising from human error or misconduct. Although gender differences in retraction rates have been previously observed in various contexts, no comprehensive study has explored this issue across all fields of science. This study examines gender disparities in scientific misconduct or errors, specifically focusing on differences in retraction rates between male and female first authors in relation to their research productivity. Using a dataset comprising 11,622 retracted articles and 19,475,437 non-retracted articles from the Web of Science and Retraction Watch, we investigate gender differences in retraction rates from the perspectives of retraction reasons, subject fields, and countries. Our findings indicate that male first authors have higher retraction rates, particularly for scientific misconduct such as plagiarism, authorship disputes, ethical issues, duplication, and fabrication/falsification. No significant gender differences were found in retractions attributed to mistakes. Furthermore, male first authors experience significantly higher retraction rates in biomedical and health sciences, as well as in life and earth sciences, whereas female first authors have higher retraction rates in mathematics and computer science. Similar patterns are observed for corresponding authors. Understanding these gendered patterns of retraction may contribute to strategies aimed at reducing their prevalence.

cs.DL

Social media uptake of scientific journals: A comparison between X and WeChat

This study examines the social media uptake of scientific journals on two different platforms - X and WeChat - by comparing the adoption of X among journals indexed in the Science Citation Index-Expanded (SCIE) with the adoption of WeChat among journals indexed in the Chinese Science Citation Database (CSCD). The findings reveal substantial differences in platform adoption and user engagement, shaped by local contexts. While only 22.7% of SCIE journals maintain an X account, 84.4% of CSCD journals have a WeChat official account. Journals in Life Sciences & Biomedicine lead in uptake on both platforms, whereas those in Technology and Physical Sciences show high WeChat uptake but comparatively lower presence on X. User engagement on both platforms is dominated by low-effort interactions rather than more conversational behaviors. Correlation analyses indicate weak-to-moderate relationships between bibliometric indicators and social media metrics, confirming that online engagement reflects a distinct dimension of journal impact, whether on an international or a local platform. These findings underscore the need for broader social media metric frameworks that incorporate locally dominant platforms, thereby offering a more comprehensive understanding of science communication practices across diverse social media and contexts.

cs.DL

Can news and social media attention reduce the influence of problematic research?

News and social media are widely used to disseminate science, but do they also help raise awareness of problems in research? This study investigates whether high levels of news and social media attention might accelerate the retraction process and increase the visibility of retracted articles. To explore this, we analyzed 15,642 news mentions, 6,588 blog mentions, and 404,082 X mentions related to 15,461 retracted articles. Articles receiving high levels of news and X mentions were retracted more quickly than non-mentioned articles in the same broad field and with comparable publication years, author impact, and journal impact. However, this effect was not statistically signicant for articles with high levels of blog mentions. Notably, articles frequently mentioned in the news experienced a significant increase in annual citation rates after their retraction, possibly because media exposure enhances the visibility of retracted articles, making them more likely to be cited. These findings suggest that increased public scrutiny can improve the efficiency of scientific self-correction, although mitigating the influence of retracted articles remains a gradual process.

cs.DL

Science cited in policy documents: Evidence from the Overton database

To reflect the extent to which science is cited in policy documents, this paper explores the presence of policy document citations for over 18 million Web of Science-indexed publications published between 2010 and 2019. Enabled by the policy document citation data provided by Overton, a searchable index of policy documents worldwide, the results show that there are 3.9% of publications in the dataset cited at least once by policy documents. Policy document citations present a citation delay towards newly published publications and show a stronger predominance to the document types of review and article. Based on the Overton database, publications in the field of Social Sciences and Humanities have the highest relative presence in policy document citations, followed by Life and Earth Sciences and Biomedical and Health Sciences. Our findings shed light not only on the impact of scientific knowledge on the policy-making process, but also on the particular focus of policy documents indexed by Overton on specific research areas.

cs.DL

A multi-dimensional analysis of usage counts, Mendeley readership, and citations for journal and conference papers

This study analyzed 16,799 journal papers and 98,773 conference papers published by IEEE Xplore in 2016 to investigate the relationships among usage counts, Mendeley readership, and citations through descriptive, regression, and mediation analyses. Differences in the relationship among these metrics between journal and conference papers are also studied. Results showed that there is no significant difference between journal and conference papers in the distribution patterns and accumulation rates of the three metrics. However, the correlation coefficients of the interrelationships between the three metrics were lower in conference papers compared to journal papers. Secondly, funding, international collaboration, and open access are positively associated with all three metrics, except for the case of funding on the usage metrics of conference papers. Furthermore, early Mendeley readership is a better predictor of citations than early usage counts and performs better for journal papers. Finally, we reveal that early Mendeley readership partially mediates between early usage counts and citation counts in the journal and conference papers. The main difference is that conference papers rely more on the direct effect of early usage counts on citations. This study contributes to expanding the existing knowledge on the relationships among usage counts, Mendeley readership, and citations in journal and conference papers, providing new insights into the relationship between the three metrics through mediation analysis.

cs.DL

How are exclusively data journals indexed in major scholarly databases? An examination of the Web of Science, Scopus, Dimensions, and OpenAlex

As part of the data-driven paradigm and open science movement, the data paper is becoming a popular way for researchers to publish their research data, based on academic norms that cross knowledge domains. Data journals have also been created to host this new academic genre. The growing number of data papers and journals has made them an important large-scale data source for understanding how research data is published and reused in our research system. One barrier to this research agenda is a lack of knowledge as to how data journals and their publications are indexed in the scholarly databases used for quantitative analysis. To address this gap, this study examines how a list of 18 exclusively data journals (i.e., journals that primarily accept data papers) are indexed in four popular scholarly databases: the Web of Science, Scopus, Dimensions, and OpenAlex. We investigate how comprehensively these databases cover the selected data journals and, in particular, how they present the document type information of data papers. We find that the coverage of data papers, as well as their document type information, is highly inconsistent across databases, which creates major challenges for future efforts to study them quantitatively. As a result, we argue that efforts should be made by data journals and databases to improve the quality of metadata for this emerging genre.

cs.DL

WeChat uptake of Chinese scholarly journals: an analysis of CSSCI-indexed journals

The study of how science is discussed and how scholarly actors interact on social media has increasingly become popular in the field of scientometrics in recent years. While most prior studies focused on research outputs discussed on global platforms, such as Twitter or Facebook, the presence of scholarly journals on local platforms was seldom studied, especially in the Chinese social media context. To fill this gap, this study investigates the uptake of WeChat (a Chinese social network app) by the Chinese scholarly journals indexed by the Chinese Social Sciences Citation Index (CSSCI). The results show that 65.3% of CSSCI-indexed journals have created WeChat public accounts and posted over 193 thousand WeChat posts in total. At the journal level, bibliometric indicators (e.g., citations, downloads, and journal impact factors) and WeChat indicators (e.g., clicks, likes, replies, and recommendations) are weakly correlated with each other, reinforcing the idea of fundamentally differentiated dimensions of indicators between bibliometrics and social media metrics. Results also show that journals with WeChat public accounts slightly outperform those without WeChat public accounts in terms of citation impact, suggesting that the WeChat presence of scientific journals is mostly positively associated with their citation impact.

cs.DL

Studying the scientific mobility and international collaboration funded by the China Scholarship Council

Every year many scholars are funded by the China Scholarship Council (CSC). The CSC is a funding agency established by the Chinese government with the main initiative of training Chinese scholars to conduct research abroad and to promote international collaboration. In this study, we identified these CSC-funded scholars sponsored by the China Scholarship Council based on the acknowledgments text indexed by the Web of Science. Bibliometric data of their publications were collected to track their scientific mobility in different fields, and to evaluate the performance of the CSC scholarship in promoting international collaboration by sponsoring the mobility of scholars. Papers funded by the China Scholarship Council are mainly from the fields of natural sciences and engineering sciences. There are few CSC-funded papers in the field of social sciences and humanities. CSC-funded scholars from mainland China have the United States, Australia, Canada, and some European countries, such as Germany, the UK, and the Netherlands, as their preferential mobility destinations across all fields of science. CSC-funded scholars published most of their papers with international collaboration during the mobility period, with a decrease in the share of international collaboration after the support of the scholarship.

cs.DL

Data sharing practices across knowledge domains: a dynamic examination of data availability statements in PLOS ONE publications

As the importance of research data gradually grows in sciences, data sharing has come to be encouraged and even mandated by journals and funders in recent years. Following this trend, the data availability statement has been increasingly embraced by academic communities as a means of sharing research data as part of research articles. This paper presents a quantitative study of which mechanisms and repositories are used to share research data in PLOS ONE articles. We offer a dynamic examination of this topic from the disciplinary and temporal perspectives based on all statements in English-language research articles published between 2014 and 2020 in the journal. We find a slow yet steady growth in the use of data repositories to share data over time, as opposed to sharing data in the paper or supplementary materials; this indicates improved compliance with the journal's data sharing policies. We also find that multidisciplinary data repositories have been increasingly used over time, whereas some disciplinary repositories show a decreasing trend. Our findings can help academic publishers and funders to improve their data sharing policies and serve as an important baseline dataset for future studies on data sharing activities.

cs.DL

Finite Volume Element Methods for Two-Dimensional Time Fractional Reaction-Diffusion Equations on Triangular Grids

In this paper, the time fractional reaction-diffusion equations with the Caputo fractional derivative are solved by using the classical $L1$-formula and the finite volume element (FVE) methods on triangular grids. The existence and uniqueness for the fully discrete FVE scheme are given. The stability result and optimal \textit{a priori} error estimate in $L^2(Ω)$-norm are derived, but it is difficult to obtain the corresponding results in $H^1(Ω)$-norm, so another analysis technique is introduced and used to achieve our goal. Finally, two numerical examples in different spatial dimensions are given to verify the feasibility and effectiveness.

math.NA

An extensive analysis of the presence of altmetric data for Web of Science publications across subject fields and research topics

Sufficient data presence is one of the key preconditions for applying metrics in practice. Based on both Altmetric.com data and Mendeley data collected up to 2019, this paper presents a state-of-the-art analysis of the presence of 12 kinds of altmetric events for nearly 12.3 million Web of Science publications published between 2012 and 2018. Results show that even though an upward trend of data presence can be observed over time, except for Mendeley readers and Twitter mentions, the overall presence of most altmetric data is still low. The majority of altmetric events go to publications in the fields of Biomedical and Health Sciences, Social Sciences and Humanities, and Life and Earth Sciences. As to research topics, the level of attention received by research topics varies across altmetric data, and specific altmetric data show different preferences for research topics, on the basis of which a framework for identifying hot research topics is proposed and applied to detect research topics with higher levels of attention garnered on certain altmetric data source. Twitter mentions and policy document citations were selected as two examples to identify hot research topics of interest of Twitter users and policy-makers, respectively, shedding light on the potential of altmetric data in monitoring research trends of specific social attention.

cs.DL