SearcharxivSearch

arXiv subjects

Mingze Zhang

Publications and source records attributed to Mingze Zhang.

4 recordsLinked to original sources

Co-leading Teams Drive Scientific Novelty in Large-scale Research Infrastructures

Large-scale research infrastructures (LSRIs) have become the engine of modern scientific discovery. While these big machines predominantly operate under a user-oriented model where external teams conduct research with support from in-house researchers, the structural integration of staff scientists into user teams and its association with scientific novelty remains unclear. By leveraging a dataset of 273,109 publications across 76 global LSRIs and applying a hybrid machine-learning framework to classify papers into three collaboration patterns: external user only, staff participating, and staff co-leading, we find a distinct novelty premium for external teams that formally integrate staff as co-authors, especially when staff scientists play co-leading rather than participating roles. Further, the premium peaks at a relatively balanced user-staff team composition, potentially due to an "epistemic lock-in" by either party. Crucially, we find that the ideal collaboration architecture evolves with user experience: while newcomers can obtain a large novelty premium from mere staff participation, experienced users only benefit from staff co-leading teams. This result suggests a "knowledge saturation effect" for which a deeper intellectual partnership is needed to sustain novelty. By revealing how user-staff collaboration structure drives scientific creativity, our study offers practical policy implications for the strategic management and intervention of LSRIs in the era of human-machine collaboration.

cs.DL

Smaller, Younger, and More Impactful: How AI-Assisted Writing Transforms Research Teams

The era of Big Science has long been defined by increasingly large and specialized research teams pushing the frontiers of knowledge. However, recent advances in artificial intelligence (AI), particularly large language models (LLMs), are beginning to reshape academic writing and scientific research, potentially disrupting the longstanding trend toward ever-larger teams and transforming other dimensions of research team structure. Drawing on 147,074 full-text publications from the PLoS family and the Nature portfolio since 2020, we examined whether and how AI-assisted writing influences team structure and team outcomes in science. Using multiple methods, including ordinary least square, quantile regression, Poisson regression, logistic regression and propensity score matching, we found that research teams using AI-assisted writing tend to be younger and smaller. Importantly, this shift toward more compact, junior-leaning teams does not come at the expense of scientific impact. On the contrary, we observed a higher probability of research teams that employed AI-assisted writing producing highly impactful publications. These results highlight the significant role of AI-assisted writing in reshaping not only how research is produced, but also how research teams are formed and assembled. Our findings call for policy improvements in research evaluation, funding, and training to address this emerging trend.

cs.CY

Scientific tools and Innovation: Big Science Facilities Yield More Novel and Interdisciplinary Knowledge

Scientific tools dictate the boundaries of human knowledge, serving as the foundation for perceptions and explorations. In the era of Big Science, science are increasingly dependent on advanced analytical technologies and experimental platforms. Over the past decades, national and supranational entities have invested massive financial resources, collaborative networks, and collective intelligence to construct Big Science Facilities (BSFs) aimed at generating cutting edge knowledge. However, empirical evaluations of these machines actual performance in driving scientific innovation remain scarce. To address this gap, we collected 310,086 publications from 88 global BSFs and constructed a matched control dataset of approximately 3 million publications sharing the same last authors. Our analysis reveals that the utilization of BSFs has expanded significantly since 1950s. Crucially, publications supported by these facilities exhibit higher recombinant novelty and interdisciplinary integration. Furthermore, this improvement is most pronounced in non physical sciences domains traditionally peripheral to BSFs core focus indicating the emergence of a powerful intra facility knowledge spillover effect. By enriching the Facilitymetrics framework, our findings provide empirical evidence that BSFs act as vital engines for scientific discovery, offering policymakers essential metrics to justify infrastructural investments, while prompting the science of science community to reassess the profound impact of scientific tools on knowledge production

cs.DL

MTMD: Multi-Scale Temporal Memory Learning and Efficient Debiasing Framework for Stock Trend Forecasting

The endeavor of stock trend forecasting is principally focused on predicting the future trajectory of the stock market, utilizing either manual or technical methodologies to optimize profitability. Recent advancements in machine learning technologies have showcased their efficacy in discerning authentic profit signals within the realm of stock trend forecasting, predominantly employing temporal data derived from historical stock price patterns. Nevertheless, the inherently volatile and dynamic characteristics of the stock market render the learning and capture of multi-scale temporal dependencies and stable trading opportunities a formidable challenge. This predicament is primarily attributed to the difficulty in distinguishing real profit signal patterns amidst a plethora of mixed, noisy data. In response to these complexities, we propose a Multi-Scale Temporal Memory Learning and Efficient Debiasing (MTMD) model. This innovative approach encompasses the creation of a learnable embedding coupled with external attention, serving as a memory module through self-similarity. It aims to mitigate noise interference and bolster temporal consistency within the model. The MTMD model adeptly amalgamates comprehensive local data at each timestamp while concurrently focusing on salient historical patterns on a global scale. Furthermore, the incorporation of a graph network, tailored to assimilate global and local information, facilitates the adaptive fusion of heterogeneous multi-scale data. Rigorous ablation studies and experimental evaluations affirm that the MTMD model surpasses contemporary state-of-the-art methodologies by a substantial margin in benchmark datasets. The source code can be found at https://github.com/MingjieWang0606/MDMT-Public.

cs.CE