SearcharxivSearch

arXiv subjects

Yanqing Hu

Publications and source records attributed to Yanqing Hu.

At least 19 recordsLinked to original sources

Notes2Skills: From Lab Notebooks to Certainty-Aware Scientific Agent Skills

Scientific discovery workflows usually contain and rely heavily on lab notes, where researchers record observations, interpret uncertain results, and plan follow-up experiments. Such informative lab notes preserve evolving scientific reasoning and author uncertainty, rather than polished final results exhibited in publications, providing a valuable opportunity for AI to engage in scientific exploration at a more comprehensive and deeper level. However, most prior work on scientific text focuses on papers, protocols, or structured databases, leaving informal laboratory notes underexplored as inputs to AI agents for science. This gap matters because lab notes often intermingle validated observations, tentative judgments, and possible experimental next steps within the same passage. If these signals are conflated, an AI agent may mistake uncertain scientific judgments for confirmed conclusions or executable actions. To this end, we present Notes2Skills, a two-stage framework for turning lab notebooks into verifiable skills for scientific AI agents while preserving the author's certainty. Across seven conditions and three wet-lab sessions, Notes2Skills is the only configuration that neither mistakes uncertain notes for firm instructions nor discards firm ones. We show that certainty preservation is the missing piece between lab notebooks and reliable agent skills, opening a path toward safer AI co-scientist systems.

cs.CL

ReviewGuard: Aligning LLM-Assisted Peer Review with Long-Term Scientific Impact

Peer review is central to scientific quality control, yet it can undervalue papers that later achieve substantial citation impact. While frontier large language models have shown promise in automating aspects of peer review, they primarily mimic human reviewer preferences rather than predict long-term scientific value. We introduce ReviewGuard, a two-stage framework that aligns LLM-generated reviews with citation-based estimates of long-term scientific impact rather than contemporaneous reviewer judgments. On 20,861 AI/ML papers from OpenReview augmented with Semantic Scholar citation data, ReviewGuard achieves a Spearman correlation of \r{ho} = 0.776 with future citations on rejected-then-published papers, outperforming human reviewers (\r{ho} = 0.492) and a supervised Expert model (\r{ho} = 0.681). Under the same decision threshold, ReviewGuard flags 10.2% of high-impact rejected papers, compared with 1.8% for human reviewers, corresponding to a 5.6x improvement. Our results demonstrate that impact-aligned reinforcement learning can provide editors with a complementary signal for identifying high-potential work, without replacing human judgment.

cs.DL

Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter

Large language model (LLM) unlearning aims to remove specific data influences from pre-trained model without costly retraining, addressing privacy, copyright, and safety concerns. However, recent studies reveal a critical vulnerability: unlearned models rapidly recover "forgotten" knowledge through relearning attacks. This fragility raises serious security concerns, especially for open-weight models. In this work, we investigate the fundamental mechanism underlying this fragility from a representation geometry perspective. We discover that existing unlearning methods predominantly optimize along dominant components, leaving minor components largely unchanged. Critically, during relearning attacks, the modifications in these dominant components are easily reversed, enabling rapid knowledge recovery, whereas minor components exhibit stronger resistance to such reversal. We further provide a theoretical analysis that explains both observations from the spectral structure of representations. Building on this insight, we propose Minor Component Unlearning (MCU), a novel unlearning approach that explicitly targets minor components in representations. By concentrating unlearning effects in these inherently robust directions, our method achieves substantially improved resistance to relearning attacks. Extensive experiments on three datasets validate our approach, demonstrating significant improvements over state-of-the-art methods including sharpness-aware minimization.

cs.CL

Restoring Network Evolution from Static Structure

The dynamical evolution of complex networks underpins the structure-function relationships in natural and artificial systems. Yet, restoring a network's formation from a single static snapshot remains challenging. Here, we present a transferable machine learning framework that infers network evolutionary trajectories solely from present topology. By integrating graph neural networks with transformers, our approach unlocks a latent temporal dimension directly from the static topology. Evaluated across diverse domains, the framework achieves high transfer accuracy of up to 95.3%, demonstrating its robustness and transferability. Applied to the Drosophila brain connectome, it restores the formation times of over 2.6 million neural connections, revealing that early-forming links support essential behaviors such as mating and foraging, whereas later-forming connections underpin complex sensory and social functions. These results demonstrate that a substantial fraction of evolutionary information is encoded within static network architecture, offering a powerful, general tool for elucidating the hidden temporal dynamics of complex systems.

physics.soc-ph

Critical Thresholds in Non-Pharmaceutical Interventions for Epidemic Control

Non-pharmaceutical interventions, such as contact tracing and social distancing, are critical for controlling epidemic outbreaks, yet their dynamic interactions remain underexplored. We introduce a probabilistic framework to analyze the synergy between contact tracing speed, quantified by the contact tracing period $\tau$, and the average number of close contacts, $\bar{k}_+$, reflecting social distancing measures. We identify critical thresholds ($R=1$) that separate pandemic and contained phases in the $\bar{k}_{+}-\tau$ plane, validated using high-resolution data from Shenzhen's 2022 Omicron outbreak (1,187 cases, 86,451 contacts). Our findings show that contact tracing alone can contain diseases with $R_0 < 2.12$ (95% CI 2.07-2.16), covering 43.33% of major infectious diseases, while combining with social distancing extends control to $R_0 < 7.82$ (95% CI 7.70-7.93), encompassing 86.67% of pathogens. These results, supported by empirical data, highlight the efficacy of rapid tracing and targeted social distancing as alternatives to mass PCR testing. Our framework offers actionable insights for optimizing NPI strategies, though challenges in scaling to regions with higher tracing miss rates or weaker infrastructure underscore the need for adaptive, data-driven policies.

physics.soc-ph

Predictability of Complex Systems

The study of complex systems has attracted widespread attention from researchers in the fields of natural sciences, social sciences, and engineering. Prediction is one of the central issues in this field. Although most related studies have focused on prediction methods, research on the predictability of complex systems has received increasing attention across disciplines--aiming to provide theories and tools to address a key question: What are the limits of prediction accuracy? Predictability itself can serve as an important feature for characterizing complex systems, and accurate estimation of predictability can provide a benchmark for the study of prediction algorithms. This allows researchers to clearly identify the gap between current prediction accuracy and theoretical limits, thereby helping them determine whether there is still significant room to improve existing algorithms. More importantly, investigating predictability often requires the development of new theories and methods, which can further inspire the design of more effective algorithms. Over the past few decades, this field has undergone significant evolution. In particular, the rapid development of data science has introduced a wealth of data-driven approaches for understanding and quantifying predictability. This review summarizes representative achievements, integrating both data-driven and mechanistic perspectives. After a brief introduction to the significance of the topic in focus, we will explore three core aspects: the predictability of time series, the predictability of network structures, and the predictability of dynamical processes. Finally, we will provide extensive application examples across various fields and outline open challenges for future research.

physics.soc-ph

Predicting the critical behavior of complex dynamic systems via learning the governing mechanisms

Critical points separate distinct dynamical regimes of complex systems, often delimiting functional or macroscopic phases in which the system operates. However, the long-term prediction of critical regimes and behaviors is challenging given the narrow set of parameters from which they emerge. Here, we propose a framework to learn the rules that govern the dynamic processes of a system. The learned governing rules further refine and guide the representative learning of neural networks from a series of dynamic graphs. This combination enables knowledge-based prediction for the critical behaviors of dynamical networked systems. We evaluate the performance of our framework in predicting two typical critical behaviors in spreading dynamics on various synthetic and real-world networks. Our results show that governing rules can be learned effectively and significantly improve prediction accuracy. Our framework demonstrates a scenario for facilitating the representability of deep neural networks through learning the underlying mechanism, which aims to steer applications for predicting complex behavior that learnable physical rules can drive.

physics.soc-ph

Spreading dynamics of information on online social networks

Social media is profoundly changing our society with its unprecedented spreading power. Due to the complexity of human behaviors and the diversity of massive messages, the information spreading dynamics are complicated, and the reported mechanisms are different and even controversial. Based on data from mainstream social media platforms, including WeChat, Weibo, and Twitter, cumulatively encompassing a total of 7.45 billion users, we uncover a ubiquitous mechanism that the information spreading dynamics are basically driven by the interplay of social reinforcement and social weakening effects. Accordingly, we propose a concise equation, which, surprisingly, can well describe all the empirical large-scale spreading trajectories. Our theory resolves a number of controversial claims and satisfactorily explains many phenomena previously observed. It also reveals that the highly clustered nature of social networks can lead to rapid and high-frequency information bursts with relatively small coverage per burst. This vital feature enables social media to have a high capacity and diversity for information dissemination, beneficial for its ecological development.

physics.soc-ph

Reconstructing the evolution history of networked complex systems

The evolution processes of complex systems carry key information in the systems' functional properties. Applying machine learning algorithms, we demonstrate that the historical formation process of various networked complex systems can be extracted, including protein-protein interaction, ecology, and social network systems. The recovered evolution process has demonstrations of immense scientific values, such as interpreting the evolution of protein-protein interaction network, facilitating structure prediction, and particularly revealing the key co-evolution features of network structures such as preferential attachment, community structure, local clustering, degree-degree correlation that could not be explained collectively by previous theories. Intriguingly, we discover that for large networks, if the performance of the machine learning model is slightly better than a random guess on the pairwise order of links, reliable restoration of the overall network formation process can be achieved. This suggests that evolution history restoration is generally highly feasible on empirical networks.

physics.soc-ph

Random node reinforcement and $K$-core structure of complex networks

To enhance robustness of complex networked systems, a simple method is introducing reinforced nodes which always function during failure propagation. A random scheme of node reinforcement can be considered as a benchmark for finding an optimal reinforcement solution. Yet there still lacks a systematic evaluation on how node reinforcement affects network structure at a mesoscopic level upon failures. Here we study this problem through the lens of $K$-cores of networks. Based on an analytical percolation framework, we first show that, on uncorrelated random graphs, with a critical size of reinforced nodes, an abrupt emergence of $K$-cores is smoothed out to a continuous one, and a detailed phase diagram is derived. We then show that, with a cost-benefit analysis on random reinforcement, for proper weight factors in cost functions with constant and increasing marginal costs, a gain function shows a unimodality, thus we can analytically find an optimal reinforcement fraction by locating the maximal gain. In all, our framework offers a gain-oriented analytical perspective to designing robust interconnected systems.

physics.soc-ph

Identify Hidden Spreaders of Pandemic over Contact Tracing Networks

The COVID-19 infection cases have surged globally, causing devastations to both the society and economy. A key factor contributing to the sustained spreading is the presence of a large number of asymptomatic or hidden spreaders, who mix among the susceptible population without being detected or quarantined. Here we propose an effective non-pharmacological intervention method of detecting the asymptomatic spreaders in contact-tracing networks, and validated it on the empirical COVID-19 spreading network in Singapore. We find that using pure physical spreading equations, the hidden spreaders of COVID-19 can be identified with remarkable accuracy. Specifically, based on the unique characteristics of COVID-19 spreading dynamics, we propose a computational framework capturing the transition probabilities among different infectious states in a network, and extend it to an efficient algorithm to identify asymptotic individuals. Our simulation results indicate that a screening method using our prediction outperforms machine learning algorithms, e.g. graph neural networks, that are designed as baselines in this work, as well as random screening of infection's closest contacts widely used by China in its early outbreak. Furthermore, our method provides high precision even with incomplete information of the contract-tracing networks. Our work can be of critical importance to the non-pharmacological interventions of COVID-19, especially with increasing adoptions of contact tracing measures using various new technologies. Beyond COVID-19, our framework can be useful for other epidemic diseases that also feature asymptomatic spreading

physics.soc-ph

Detecting and modelling real percolation and phase transitions of information on social media

It is widely believed that information spread on social media is a percolation process, with parallels to phase transitions in theoretical physics. However, evidence for this hypothesis is limited, as phase transitions have not been directly observed in any social media. Here, through analysis of 100 million Weibo and 40 million Twitter users, we identify percolation-like spread, and find that it happens more readily than current theoretical models would predict. The lower percolation threshold can be explained by the existence of positive feedback in the coevolution between network structure and user activity level, such that more active users gain more followers. Moreover, this coevolution induces an extreme imbalance in users' influence. Our findings indicate that the ability of information to spread across social networks is higher than expected, with implications for many information spread problems.

physics.soc-ph

Induced Percolation on Networked Systems

Percolation theory has been widely used to study phase transitions in complex networked systems. It has also successfully explained several macroscopic phenomena across different fields. Yet, the existent theoretical framework for percolation places the focus on the direct interactions among the system's components, while recent empirical observations have shown that indirect interactions are common in many systems like ecological and social networks, among others. Here, we propose a new percolation framework that accounts for indirect interactions, which allows to generalize the current theoretical body and understand the role of the underlying indirect influence of the components of a networked system on its macroscopic behavior. We report a rich phenomenology in which first-order, second-order or hybrid phase transitions are possible depending on whether the links of the substrate network are directed, undirected or a mix, respectively. We also present an analytical framework to characterize the proposed induced percolation, paving the way to further understand network dynamics with indirect interactions.

physics.soc-ph

A Novel Framework with Information Fusion and Neighborhood Enhancement for User Identity Linkage

User identity linkage across social networks is an essential problem for cross-network data mining. Since network structure, profile and content information describe different aspects of users, it is critical to learn effective user representations that integrate heterogeneous information. This paper proposes a novel framework with INformation FUsion and Neighborhood Enhancement (INFUNE) for user identity linkage. The information fusion component adopts a group of encoders and decoders to fuse heterogeneous information and generate discriminative node embeddings for preliminary matching. Then, these embeddings are fed to the neighborhood enhancement component, a novel graph neural network, to produce adaptive neighborhood embeddings that reflect the overlapping degree of neighborhoods of varying candidate user pairs. The importance of node embeddings and neighborhood embeddings are weighted for final prediction. The proposed method is evaluated on real-world social network data. The experimental results show that INFUNE significantly outperforms existing state-of-the-art methods.

cs.SI

Beyond the Coverage of Information Spreading: Analytical and Empirical Evidence of Re-exposure in Large-scale Online Social Networks

Peer influence and social contagion are key denominators in the adoption and participation of information spreading, such as news propagation, word-of-mouth or viral marketing. In this study, we argue that it is biased to only focus on the scale and coverage of information spreading, and propose that the level of influence reinforcement, quantified by the re-exposure rate, i.e., the rate of individuals who are repeatedly exposed to the same information, should be considered together to measure the effectiveness of spreading. We show that local network structural characteristics significantly affects the probability of being exposed or re-exposed to the same information. After analyzing trending news on the super large-scale online network of Sina Weibo (China's Twitter) with 430 million connected users, we find a class of users with extremely low exposure rate, even they are following tens of thousands of others; and the re-exposure rate is substantially higher for news with more transmission waves and stronger secondary forwarding. While exposure and re-exposure rate typically grow together with the scale of spreading, we find exceptional cases where it is possible to achieve a high exposure rate while maintaining low re-exposure rate, or vice versa.

physics.soc-ph

Coevolution spreading in complex networks

The propagations of diseases, behaviors and information in real systems are rarely independent of each other, but they are coevolving with strong interactions. To uncover the dynamical mechanisms, the evolving spatiotemporal patterns and critical phenomena of networked coevolution spreading are extremely important, which provide theoretical foundations for us to control epidemic spreading, predict collective behaviors in social systems, and so on. The coevolution spreading dynamics in complex networks has thus attracted much attention in many disciplines. In this review, we introduce recent progress in the study of coevolution spreading dynamics, emphasizing the contributions from the perspectives of statistical mechanics and network science. The theoretical methods, critical phenomena, phase transitions, interacting mechanisms, and effects of network topology for four representative types of coevolution spreading mechanisms, including the coevolution of biological contagions, social contagions, epidemic-awareness, and epidemic-resources, are presented in detail, and the challenges in this field as well as open issues for future studies are also discussed.

physics.soc-ph

Local structure can identify and quantify influential global spreaders in large scale social networks

Measuring and optimizing the influence of nodes in big-data online social networks are important for many practical applications, such as the viral marketing and the adoption of new products. As the viral spreading on social network is a global process, it is commonly believed that measuring the influence of nodes inevitably requires the knowledge of the entire network. Employing percolation theory, we show that the spreading process displays a nucleation behavior: once a piece of information spread from the seeds to more than a small characteristic number of nodes, it reaches a point of no return and will quickly reach the percolation cluster, regardless of the entire network structure, otherwise the spreading will be contained locally. Thus, we find that, without the knowledge of entire network, any nodes' global influence can be accurately measured using this characteristic number, which is independent of the network size. This motivates an efficient algorithm with constant time complexity on the long standing problem of best seed spreaders selection, with performance remarkably close to the true optimum.

physics.soc-ph

Non-trivial Resource Amount Requirement in the Early Stage for Containing Fatal Diseases

During an epidemic control, the containment of the disease is usually achieved through increasing devoted resource to shorten the duration of infectiousness. However, the impact of this resource expenditure has not been studied quantitatively. Using the well-documented cholera data, we observe empirically that the recovery rate which is related to the duration of infectiousness has a strong positive correlation with the average resource devoted to the infected individuals. By incorporating this relation we build a novel model and find that insufficient resource leads to an abrupt increase in the infected population size, which is in marked contrast with the continuous phase transitions believed previously. Counterintuitively, this abrupt phase transition is more pronounced in the less contagious diseases, which usually correspond to the most fatal ones. Furthermore, we find that even for a single infection source, public resource needs to meet a significant amount, which is proportional to the whole population size to ensure epidemic containment. Our findings provide a theoretical foundation for efficient epidemic containment strategies in the early stage.

physics.soc-ph