SearcharxivSearch

arXiv subjects

Jin Mao

Publications and source records attributed to Jin Mao.

8 recordsLinked to original sources

From citation intent to knowledge contribution: Classifying what cited papers actually contribute

Understanding the flow and evolution of scientific knowledge is essential for assessing research impact. Existing citation analysis methods mainly focus on citing authors' subjective intents, failing to consistently characterize cited papers' knowledge contributions. This study proposes the Knowledge Contribution Taxonomy (KCT), derived from the Scientific Research Logic Model, which identifies the type of knowledge a cited paper contributes based on the citation context. KCT classifies citations into Method, Resource Tool, Empirical Finding, and Background, further distinguishing core from non-core contributions. We propose a Dual-Path Fusion model for the classification task, which achieves an accuracy of 85.5%, outperforming mainstream large language models. An analysis of 802,202 citations from the ACL Anthology reveals that core knowledge contributions account for only 39.09% of all citations. The core knowledge contribution citation count achieves higher hit rates for award-winning papers than the traditional citation count at all ranking cutoffs, reflecting the value of differentiating knowledge contributions for research evaluation and impact prediction. In dissemination prediction experiments, KCT outperforms citation intent classification, demonstrating its stronger predictive validity for scholarly dissemination. By focusing on the knowledge contributions of cited papers, the KCT can support differentiated research evaluation.

cs.DL

A large-scale dataset of sub-institution name disambiguation and hierarchical structures from OpenAlex

Accurate attribution of scholarly work to specific sub-institutional units, such as schools or departments of a university, is crucial for granular research assessment and policymaking. While robust identifiers exist for top-level institutions, standardized data for sub-level units remains scarce due to the linguistic and structural variability of affiliation strings. In this study, we introduce OpenSubAffil, a large-scale dataset mapping raw affiliation strings from OpenAlex to disambiguated sub-institutional entities and their hierarchical structures. We developed a pipeline integrating named entity recognition (NER) with embedding-based clustering. Furthermore, we proposed a multi-signal scoring function that synthesizes lexical and co-occurrence evidence to reconstruct the sub-institutional hierarchy. OpenSubAffil comprises mappings for 40 million affiliation strings to 638,843 disambiguated sub-units across 18,635 top-level institutions, together with their hierarchical relationships. Validation against Wikidata benchmarks and manual investigation show that our method achieves promising performance. Overall, this dataset bridges the granularity gap between individual researchers and top-level institutions, enabling high-resolution analyses of scholarly output and communication at the sub-institutional level. The OpenSubAffil dataset is publicly available at https://doi.org/10.5281/zenodo.19602782.

cs.DL

Citation importance-aware document representation learning for large-scale science mapping

Effective science mapping relies on high-quality representations of scientific documents. As an important task in scientometrics and information studies, science mapping is often challenged by the complex and heterogeneous nature of citations. While previous studies have attempted to improve document representations by integrating citation and semantic information, the heterogeneity of citations is often overlooked. To address this problem, this study proposes a citation importance-aware contrastive learning framework that refines the supervisory signal. We first develop a scalable measurement of citation importance based on location, frequency, and self-citation characteristics. Citation importance is then integrated into the contrastive learning process through an importance-aware sampling strategy, which selects low-importance citations as hard negatives. This forces the model to learn finer-grained representations that distinguish between important and perfunctory citations. To validate the effectiveness of the proposed framework, we fine-tune a SciBERT model and perform extensive evaluations on SciDocs and PubMed benchmark datasets. Results show consistent improvements in both document representation quality and science mapping accuracy. Furthermore, we apply the trained model to over 33 million documents from Web of Science. The resulting map of science accurately visualizes the global and local intellectual structure of science and reveals interdisciplinary research fronts. By operationalizing citation heterogeneity into a scalable computational framework, this study demonstrates how differentiating citations by their importance can be effectively leveraged to improve document representation and science mapping.

cs.DL

Learned-Rule-Augmented Large Language Model Evaluators

Large language models (LLMs) are predominantly used as evaluators for natural language generation (NLG) tasks, but their application to broader evaluation scenarios remains limited. In this work, we explore the potential of LLMs as general evaluators across diverse tasks. Although LLM-based evaluators have made progress in different areas, existing methods struggle to generalize due to their reliance on costly, human-designed evaluation principles, which are often misaligned with both annotated data and LLMs' understanding.To address these challenges, we propose a rule-augmented evaluation paradigm. First, we introduce a rule distillation method that automatically extracts scoring rules from data using an LLM-assisted Monte Carlo Tree Search (MCTS), alleviating scalability issues and improving alignment with data. Second, to enable LLMs to effectively apply the learned rules, we propose two strategies: (1) Chain-of-Rule (CoR), which guides LLM to follow distilled rules, and (2) training a rule-augmented LLM evaluator (RuAE) via reinforcement learning, further bridging the gap between rules and LLMs' reasoning. Extensive experiments on diverse tasks demonstrate the effectiveness and generalizability of our approach across various evaluation scenarios.

cs.AI

The geography of novel and atypical research

The production of knowledge has become increasingly a global endeavor. Yet, location related factors, such as local working environment and national policy designs, may continue to affect what kind of science is being pursued. Here we examine the geography of the production of creative science by country, through the lens of novelty and atypicality proposed in Uzzi et al. (2013). We quantify a country's representativeness in novel and atypical science, finding persistent differences in propensity to generate creative works, even among developed countries that are large producers in science. We further cluster countries based on how their tendency to publish novel science changes over time, identifying one group of emerging countries. Our analyses point out the recent emergence of China not only as a large producer in science but also as a leader that disproportionately produces more novel and atypical research. Discipline specific analysis indicates that China's over-production of atypical science is limited to a few disciplines, especially its most prolific ones like materials science and chemistry.

physics.soc-ph

Examining Linguistic Shifts in Academic Writing Before and After the Launch of ChatGPT: A Study on Preprint Papers

Large Language Models (LLMs), such as ChatGPT, have prompted academic concerns about their impact on academic writing. Existing studies have primarily examined LLM usage in academic writing through quantitative approaches, such as word frequency statistics and probability-based analyses. However, few have systematically examined the potential impact of LLMs on the linguistic characteristics of academic writing. To address this gap, we conducted a large-scale analysis across 823,798 abstracts published in last decade from arXiv dataset. Through the linguistic analysis of features such as the frequency of LLM-preferred words, lexical complexity, syntactic complexity, cohesion, readability and sentiment, the results indicate a significant increase in the proportion of LLM-preferred words in abstracts, revealing the widespread influence of LLMs on academic writing. Additionally, we observed an increase in lexical complexity and sentiment in the abstracts, but a decrease in syntactic complexity, suggesting that LLMs introduce more new vocabulary and simplify sentence structure. However, the significant decrease in cohesion and readability indicates that abstracts have fewer connecting words and are becoming more difficult to read. Moreover, our analysis reveals that scholars with weaker English proficiency were more likely to use the LLMs for academic writing, and focused on improving the overall logic and fluency of the abstracts. Finally, at discipline level, we found that scholars in Computer Science showed more pronounced changes in writing style, while the changes in Mathematics were minimal.

cs.CL

A GAN-based Semantic Communication for Text without CSI

Recently, semantic communication (SC) has been regarded as one of the potential paradigms of 6G. Current SC frameworks require channel state information (CSI) to handle severe signal distortion induced by channel fading. Since the channel estimation overhead for obtaining CSI cannot be neglected, we therefore propose a generative adversarial network (GAN) based SC framework (Ti-GSC) that doesn't require CSI. In Ti-GSC, two main modules, i.e., an autoencoder-based encoder-decoder module (AEDM) and a GAN-based signal distortion suppression module (GSDSM) are included where AEDM first encodes the data at the source before transmission, and then GSDSM suppresses the distortion of the received signals in both syntactic and semantic dimensions at the destination. At last, AEDM decodes the distortion-suppressed signal at the destination. To measure signal distortion, syntactic distortion and semantic distortion terms are newly added to the total loss function. To achieve better training results, joint optimization-based training (JOT) and alternating optimization-based training (AOT) are designed for the proposed Ti-GSC. Experimental results show that JOT is more efficient for Ti-GSC. Moreover, without CSI, bilingual evaluation understudy (BLEU) score achieved by Ti-GSC is about 40% and 62% higher than that achieved by existing SC frameworks in Rician and Rayleigh fading, respectively. (*Due to the notification of arXiv "The Abstract field cannot be longer than 1,920 characters", the appeared Abstract is shortened. For the full Abstract, please download the Article.)

cs.IT

Finding citations for PubMed: A large-scale comparison between five freely available bibliographic data sources

As an important biomedical database, PubMed provides users with free access to abstracts of its documents. However, citations between these documents need to be collected from external data sources. Although previous studies have investigated the coverage of various data sources, the quality of citations is underexplored. In response, this study compares the coverage and citation quality of five freely available data sources on 30 million PubMed documents, including OpenCitations Index of CrossRef open DOI-to-DOI citations (COCI), Dimensions, Microsoft Academic Graph (MAG), National Institutes of Health Open Citation Collection (NIH-OCC), and Semantic Scholar Open Research Corpus (S2ORC). Three gold standards and five metrics are introduced to evaluate the correctness and completeness of citations. Our results indicate that Dimensions is the most comprehensive data source that provides references for 62.4% of PubMed documents, outperforming the official NIH-OCC dataset (56.7%). Over 90% of citation links in other data sources can also be found in Dimensions. The coverage of MAG, COCI, and S2ORC is 59.6%, 34.7%, and 23.5%, respectively. Regarding the citation quality, Dimensions and NIH-OCC achieve the best overall results. Almost all data sources have a precision higher than 90%, but their recall is much lower. All databases have better performances on recent publications than earlier ones. Meanwhile, the gaps between different data sources have diminished for the documents published in recent years. This study provides evidence for researchers to choose suitable PubMed citation sources, which is also helpful for evaluating the citation quality of free bibliographic databases.

cs.DL