SearcharxivSearch

arXiv subjects

Zhesi Shen

Publications and source records attributed to Zhesi Shen.

At least 19 recordsLinked to original sources

The Landscape of problematic papers in the field of non-coding RNA

Retractions have increased sharply in recent years, alongside a growing number of papers that receive post-publication comments questioning their reliability (commented papers). Together, retracted and commented papers undermine the credibility of scientific research and may also threaten public health. In this study, we examine problematic papers in the field of non-coding RNA (ncRNA) from multiple perspectives to identify common patterns and inform strategies for addressing large-scale fraudulent publications. We find that studies on under-investigated ncRNAs are more likely to become problematic papers. These papers often show substantial textual similarity, and many additional papers with similar text also display suspicious image duplication. Healthcare institutions, particularly those with lower publication output, appear especially vulnerable to producing such papers. Most problematic papers are concentrated in a small set of journals, many of which do not adequately address concerns raised after publication. Overall, our findings indicate that a substantial number of problematic papers may remain undetected and that their shared characteristics can support more effective strategies for identifying and curbing large-scale fraudulent publications.

cs.DL

Mapping Academic Integrity: Global Retraction Trends Explored through a Topic Lens

Scientific publications have long served as the cornerstone of innovation, exhibiting stable growth over the years. Recently, however, retractions have surged dramatically, driven largely by the proliferation of low-quality and fraudulent articles, posing a substantial threat to research integrity. By integrating annual publication and retraction data, this study employs the relative retraction rate (R3) to systematically examine disparities and evolving trends from a topical perspective. Our analysis reveals that the number of retractions has grown significantly faster than that of global publications, yielding an overall retraction rate of 0.12%. While retractions occur across all disciplines, substantial disparities exist, ranging from 0.035% in Physics to 0.34% in Computer Science. This gap widens at finer levels of granularity, reaching roughly 8.99% in Human-Computer Interaction. Moreover, unusually high R3 values frequently coincide with rapid publication growth in specific fields. We also developed Retraction Monitor, a web application for monitoring retraction dynamics across diverse fields, enabling stakeholders to visualize these trends and assess risks to research integrity. These findings provide valuable insights for identifying high-risk fields and developing tailored governance policies to strengthen research rigor and mitigate field-specific retraction risks.

cs.DL

Revisiting the field normalization approaches/practices

Field normalization plays a crucial role in scientometrics to ensure fair comparisons across different disciplines. In this paper, we revisit the effectiveness of several widely used field normalization methods. Our findings indicate that source-side normalization (as employed in SNIP) does not fully eliminate citation bias across different fields and the imbalanced paper growth rates across fields are a key factor for this phenomenon. To address the issue of skewness, logarithmic transformation has been applied. Recently, a combination of logarithmic transformation and mean-based normalization, expressed as ln(c+1)/mu, has gained popularity. However, our analysis shows that this approach does not yield satisfactory results. Instead, we find that combining logarithmic transformation (ln(c+1)) with z-score normalization provides a better alternative. Furthermore, our study suggests that the better performance is achieved when combining both source-side and target-side field normalization methods.

cs.DL

Is Journal Citation Indicator a good metric for Art & Humanities Journals currently?

Probably Not. Journal Citation Indicator (JCI) was introduced to address the limitations of traditional metrics like the Journal Impact Factor (JIF), particularly its inability to normalize citation impact across different disciplines. This study reveals that JCI faces significant challenges in field normalization for Art & Humanities journals, as evidenced by much lower correlations with a more granular, paper-level metric, CNCI-CT. A detailed analysis of Architecture journals highlights how journal-level misclassification and the interdisciplinary nature of content exacerbate these issues, leading to less reliable evaluations. We recommend improving journal classification systems or adopting paper-level normalization methods, potentially supported by advanced AI techniques, to enhance the accuracy and effectiveness of JCI for Art & Humanities disciplines.

cs.DL

Evaluating the Accuracy of the Labeling System in Web of Science for the Sustainable Development Goals

Monitoring and fostering research aligned with the Sustainable Development Goals (SDGs) is crucial for formulating evidence-based policies, identifying best practices, and promoting global collaboration. The key step is developing a labeling system to map research publications to their related SDGs. The SDGs labeling system integrated in Web of Science (WoS), which assigns citation topics instead of individual publication to SDGs, has emerged as a promising tool.However we still lack of a comprehensive evaluation of the performance of WoS labeling system. By comparing with the Bergon approach, we systematically assessed the relatedness between citation topics and SDGs. Our analysis identified 15% of topics showing low relatedness to their assigned SDGs at a 1% threshold. Notably, SDGs such as '11 Cities', '07 Energy', and '13 Climate' exhibited higher percentages of low related topics. In addition, we revealed that certain topics are significantly underrepresented in their relevant SDGs, particularly for '02 Hunger', '12 Consumption', and '15 Land'. This study underscores the critical need for continual refinement and validation of SDGs labeling systems in WoS.

cs.DL

Comparison of Sustainable Development Goals Labeling Systems based on Topic Coverage

With the growing importance of sustainable development goals (SDGs), various labeling systems have emerged for effective monitoring and evaluation. This study assesses six labeling systems across 1.85 million documents at both paper level and topic level. Our findings indicate that the SDGO and SDSN systems are more aggressive, while systems such as Auckland, Aurora, SIRIS, and Elsevier exhibit significant topic consistency, with similarity scores exceeding 0.75 for most SDGs. However, similarities at the paper level generally fall short, particularly for specific SDGs like SDG 10. We highlight the crucial role of contextual information in keyword-based labeling systems, noting that overlooking context can introduce bias in the retrieval of papers (e.g., variations in "migration" between biomedical and geographical contexts). These results reveal substantial discrepancies among SDG labeling systems, emphasizing the need for improved methodologies to enhance the accuracy and relevance of SDG evaluations.

cs.DL

The Unique Citing Documents Journal Impact Factor (Uniq-JIF) as a Supplement for the standard Journal Impact Factor

This paper introduces the Unique Citing Documents Journal Impact Factor(Uniq-JIF) as a supplement to the traditional Journal Impact Factor(JIF). The Uniq-JIF counts each citing document only once, aiming to reduce the effects of citation manipulations. Analysis of 2023 Journal Citation Reports data shows that for most journals, the Uniq-JIF is less than 20% lower than the JIF, though some journals show a drop of over 75%. The Uniq-JIF also highlights significant reductions for journals suppressed due to citation issues, indicating its effectiveness in identifying problematic journals. The Uniq-JIF offers a more nuanced view of a journal's influence and can help reveal journals needing further scrutiny.

cs.DL

An Explorative Study on Document Type Assignment of Review Articles in Web of Science, Scopus and Journals' Website

Accurately assigning the document type of review articles in citation index databases like Web of Science(WoS) and Scopus is important. This study aims to investigate the document type assignation of review articles in web of Science, Scopus and Journals' website in a large scale. 27,616 papers from 160 journals from 10 review journal series indexed in SCI are analyzed. The document types of these papers labeled on journals' website, and assigned by WoS and Scopus are retrieved and compared to determine the assigning accuracy and identify the possible reasons of wrongly assigning. For the document type labeled on the website, we further differentiate them into explicit review and implicit review based on whether the website directly indicating it is review or not. We find that WoS and Scopus performed similarly, with an average precision of about 99% and recall of about 80%. However, there were some differences between WoS and Scopus across different journal series and within the same journal series. The assigning accuracy of WoS and Scopus for implicit reviews dropped significantly. This study provides a reference for the accuracy of document type assigning of review articles in WoS and Scopus, and the identified pattern for assigning implicit reviews may be helpful to better labeling on website, WoS and Scopus.

cs.DL

Novel utilization of a paper-level classification system for the evaluation of journal impact: An update of the CAS Journal Ranking

Since its first release in 2004, the CAS Journal Ranking, a ranking system of journals based on a citation impact indicator, has been widely used both in selecting journals when submitting manuscripts and conducting research evaluation in China This paper introduces an upgraded version of the CAS Journal Ranking released in 2020 and the corresponding improvements. We will discuss the following improvements: (1) the CWTS paper-level classification system, a fine-grained classification system, utilized for field normalization, (2) the Field Normalized Citation Success Index (FNCSI), an indicator which is robust against not only extremely highly cited publications, but also wrongly assigned document types, and (3) document type difference. In addition, this paper will present part of the ranking results and an interpretation of the features of the FNCSI indicator.

cs.DL

Analyzing Journal Category Assignment Using a Paper-level Classification System: Multidisciplinary Sciences Journals

In the field of scientometrics, the subject classification system of academic journals holds great importance. Accurate identification and classification of "multidisciplinary" journals are crucial in revealing the scientific structure and evaluating journals. Based on data from the Web of Science database from 2016 to 2020, we calculated the disciplinary diversity of journals using the paper-level subject classification system, then conducted a systematic analysis of JCR multidisciplinary journals. Studies showed that most multidisciplinary journals have high disciplinary diversity, while non-multidisciplinary journals tend to have relatively lower diversity. Some multidisciplinary journals with low disciplinary diversities may misclassify disciplines. In addition, there are inconsistencies in the diversity of journal disciplines at different granularities. Our study also visually analyzed the four types of diversity distribution tendencies of multidisciplinary journals. Moreover, ten potential multidisciplinary journals were found in non-multidisciplinary categories.

cs.DL

Two indicators rule them all: Mean and standard deviation used to calculate other journal indicators based on log-normal distribution of citation counts

Two journal-level indicators, respectively the mean ($m^i$) and the standard deviation ($v^i$) are proposed to be the core indicators of each journal and we show that quite several other indicators can be calculated from those two core indicators, assuming that yearly citation counts of papers in each journal follows more or less a log-normal distribution. Those other journal-level indicators include journal h index, journal one-by-one-sample comparison citation success index $S_j^i$, journal multiple-sample $K^i-K^j$ comparison success rate $S_{j,K^j}^{i,K^i }$, and minimum representative sizes $κ_j^i$ and $κ_i^j$, the average ranking of all papers in a journal in a set of journals($R^t$). We find that those indicators are consistent with those calculated directly using the raw citation data ($C^i=\{c_1^i,c_2^i,\dots,c_{N^i}^i \},\forall i$) of journals. In addition to its theoretical significance, the ability to estimate other indicators from core indicators has practical implications. This feature enables individuals who lack access to raw citation count data to utilize other indicators by simply using core indicators, which are typically easily accessible.

cs.DL

Control core of undirected complex networks

With the development of complex networks, many researchers have paid greater attention to studying the control of complex networks over the last decade. Although some theoretical breakthroughs allow us to identify all driver nodes, we still lack an efficient method to identify the driver nodes and understand the roles of individual nodes in contributing to the control of a large complex network. Here, we apply a leaf removal process (LRP) to find a substructure of an undirected network, which is considered as the control core of the original network. Based on a strict mathematical proof, the control core obtained by the LRP has the same controllability as the original network, and it contains at least one set of driver nodes. With this method, we systematically investigate the structural property of the control core with respect to different average degrees of the original networks ($\langle k \rangle$). We denote the node density ($n_\text{core}$) and link density ($l_\text{core}$) to characterize the control core when applying the LRP, and we study the impact of $\langle k \rangle$ on $n_\text{core}$ and $l_\text{core}$ in two artificial networks: undirected Erdös-Rényi (ER) random networks and undirected scale-free (SF) networks. We find that $n_\text{core}$ and $l_\text{core}$ both change nonmonotonously with increasing $\langle k \rangle$ in the two typical undirected networks. With the aid of core percolation theory, we can offer the theoretical predictions for both $n_\text{core}$ and $l_\text{core}$ as a function of $\langle k \rangle$. Then, we recognize that finding the driver nodes in the control core is much more efficient than in the original network by comparing $n_{\text{D}}$, the controllability of the original network, and $n_{\text{core}}$, regardless of how $\langle k \rangle$ increases.

physics.soc-ph

The effect of national and international multiple affiliations on citation impact

Researchers affiliated with multiple institutions are increasingly seen in current scientific environment. In this paper we systematically analyze the multi-affiliated authorship and its effect on citation impact, with focus on the scientific output of research collaboration. By considering the nationality of each institutions, we further differentiate the national multi-affiliated authorship and international multi-affiliated authorship and reveal their different patterns across disciplines and countries. We observe a large share of publications with multi-affiliated authorship (45.6%) in research collaboration, with a larger share of publications containing national multi-affiliated authorship in medicine related and biology related disciplines, and a larger share of publications containing international type in Space Science, Physics and Geosciences. To a country-based view, we distinguish between domestic and foreign multi-affiliated authorship to a specific country. Taking G7 and BRICS countries as samples from different S&T level, we find that the domestic national multi-affiliated authorship relate to more on citation impact for most disciplines of G7 countries, while domestic international multi-affiliated authorships are more positively influential for most BRICS countries.

cs.DL

Increasing trend of scientists to switch between topics

We analyze the publication records of individual scientists, aiming to quantify the topic switching dynamics of scientists and its influence. For each scientist, the relations among her publications are characterized via shared references. We find that the co-citing network of the papers of a scientist exhibits a clear community structure where each major community represents a research topic. Our analysis suggests that scientists tend to have a narrow distribution of the number of topics. However, researchers nowadays switch more frequently between topics than those in the early days. We also find that high switching probability in early career (<12y) is associated with low overall productivity, while it is correlated with high overall productivity in latter career. Interestingly, the average citation per paper, however, is in all career stages negatively correlated with the switching probability. We propose a model with exploitation and exploration mechanisms that can explain the main observed features.

physics.soc-ph

Do Mathematicians, Economists and Biomedical Scientists Trace Large Topics More Strongly Than Physicists?

In this work, we extend our previous work on largeness tracing among physicists to other fields, namely mathematics, economics and biomedical science. Overall, the results confirm our previous discovery, indicating that scientists in all these fields trace large topics. Surprisingly, however, it seems that researchers in mathematics tend to be more likely to trace large topics than those in the other fields. We also find that on average, papers in top journals are less largeness-driven. We compare researchers from the USA, Germany, Japan and China and find that Chinese researchers exhibit consistently larger exponents, indicating that in all these fields, Chinese researchers trace large topics more strongly than others. Further correlation analyses between the degree of largeness tracing and the numbers of authors, affiliations and references per paper reveal positive correlations -- papers with more authors, affiliations or references are likely to be more largeness-driven, with several interesting and noteworthy exceptions: in economics, papers with more references are not necessary more largeness-driven, and the same is true for papers with more authors in biomedical science. We believe that these empirical discoveries may be valuable to science policy-makers.

physics.soc-ph

Locating the source of diffusion in complex networks by time-reversal backward spreading

Locating the source that triggers a dynamical process is a fundamental but challenging problem in complex networks, ranging from epidemic spreading in society and on the Internet to cancer metastasis in the human body. An accurate localization of the source is inherently limited by our ability to simultaneously access the information of all nodes in a large-scale complex network. This thus raises two critical questions: how do we locate the source from incomplete information and can we achieve full localization of sources at any possible location from a given set of observable nodes. Here we develop a time-reversal backward spreading algorithm to locate the source of a diffusion-like process efficiently and propose a general locatability condition. We test the algorithm by employing epidemic spreading and consensus dynamics as typical dynamical processes and apply it to the H1N1 pandemic in China. We find that the sources can be precisely located in arbitrary networks insofar as the locatability condition is assured. Our tools greatly improve our ability to locate the source of diffusion in complex networks based on limited accessibility of nodal information. Moreover, they have implications for controlling a variety of dynamical processes taking place on complex networks, such as inhibiting epidemics, slowing the spread of rumors, pollution control and environmental protection.

physics.soc-ph

A universal data based method for reconstructing complex networks with binary-state dynamics

To understand, predict, and control complex networked systems, a prerequisite is to reconstruct the network structure from observable data. Despite recent progress in network reconstruction, binary-state dynamics that are ubiquitous in nature, technology and society still present an outstanding challenge in this field. Here we offer a framework for reconstructing complex networks with binary-state dynamics by developing a universal data-based linearization approach that is applicable to systems with linear, nonlinear, discontinuous, or stochastic dynamics governed by monotonous functions. The linearization procedure enables us to convert the network reconstruction into a sparse signal reconstruction problem that can be resolved through convex optimization. We demonstrate generally high reconstruction accuracy for a number of complex networks associated with distinct binary-state dynamics from using binary data contaminated by noise and missing data. Our framework is completely data driven, efficient and robust, and does not require any a priori knowledge about the detailed dynamical process on the network. The framework represents a general paradigm for reconstructing, understanding, and exploiting complex networked systems with binary-state dynamics.

physics.soc-ph

Fundamental building blocks of controlling complex networks: A universal controllability framework

To understand the controllability of complex networks is a forefront problem relevant to different fields of science and engineering. Despite recent advances in network controllability theories, an outstanding issue is to understand the effect of network topology and nodal interactions on the controllability at the most fundamental level. Here we develop a universal framework based on local information only to unearth the most {\em fundamental building blocks} that determine the controllability. In particular, we introduce a network dissection process to fully unveil the origin of the role of individual nodes and links in control, giving rise to a criterion for the much needed strong structural controllability. We theoretically uncover various phase-transition phenomena associated with the role of nodes and links and strong structural controllability. Applying our theory to a large number of empirical networks demonstrates that technological networks are more strongly structurally controllable (SSC) than many social and biological networks, and real world networks are generally much more SSC than their random counterparts with intrinsic resilience and adaptability as a result of human design and natural evolution.

physics.soc-ph