Searcharxiv⌕ Search

arXiv subjects

Yongtao Zhang

Publications and source records attributed to Yongtao Zhang.

14 recordsLinked to original sources

A scalable entanglement distribution framework for hybrid quantum networks

The realization of a scalable quantum network (QN) hinges on the efficient distribution of entanglement across heterogeneous hardware platforms. While traditional QN frameworks are often restricted to single-encoding schemes, relying exclusively on either discrete (DV) or continuous (CV) variables, such segregated architectures limit the capacity and interoperability of large-scale systems. Here, we propose a scalable hybrid entanglement distribution scheme that bridges these domains by mapping hybrid entanglement swapping and concentration at DV-CV interfaces onto a unified set of series and parallel graph rules. We demonstrate that our hybrid approach exhibits a potential performance advantage over traditional point-to-point generation of maximally entangled states, achieving enhanced end-to-end entanglement within series-parallel topologies. These results provide a robust theoretical foundation for designing resource-efficient, interpretable architectures for the future quantum internet.

quant-ph↗

Negativity Percolation in Continuous-Variable Quantum Networks

Quantum networks (QNs) have been predominantly driven by discrete-variable (DV) architectures. Yet, optical platforms naturally generate Gaussian states--the common states of continuous-variable (CV) systems, making CV-based QNs an attractive route toward scalable, chip-integrated quantum computation and communication. To bridge the gap between well-studied DV entanglement percolation theories and their CV counterpart, we introduce a Gaussian-to-Gaussian entanglement distribution scheme that deterministically transports two-mode squeezed vacuum states across large CV networks. Analysis of the scheme's collective behavior using statistical-physics methods reveals a new form of entanglement percolation--negativity percolation theory (NegPT)--characterized by a bounded entanglement measure called the ratio negativity. We discover that NegPT exhibits a mixed-order phase transition, marked simultaneously by both an abrupt change in global entanglement and a long-range correlation between nodes. This distinctive behavior places CV-based QNs in a new universality class, fundamentally distinct from DV systems. Additionally, the abruptness of this transition introduces a critical vulnerability of CV-based QNs: conventional feedback mechanism becomes inherently unstable near the threshold, highlighting practical implications for stabilizing large-scale CV-based QNs. Our results unify statistical models for CV-based entanglement distribution and uncover previously unexplored critical phenomena unique to CV systems, providing valuable insights and guidelines essential for developing robust, feedback-stabilized QNs.

quant-ph↗

When to Memorize and When to Stop: Gated Recurrent Memory for Long-Context Reasoning

While reasoning over long context is crucial for various real-world applications, it remains challenging for large language models (LLMs) as they suffer from performance degradation as the context length grows. Recent work MemAgent has tried to tackle this by processing context chunk-by-chunk in an RNN-like loop and updating a textual memory for final answering. However, this naive recurrent memory update faces two crucial drawbacks: (i) memory can quickly explode because it can update indiscriminately, even on evidence-free chunks; and (ii) the loop lacks an exit mechanism, leading to unnecessary computation after even sufficient evidence is collected. To address these issues, we propose GRU-Mem, which incorporates two text-controlled gates for more stable and efficient long-context reasoning. Specifically, in GRU-Mem, the memory only updates when the update gate is open and the recurrent loop will exit immediately once the exit gate is open. To endow the model with such capabilities, we introduce two reward signals $r^{\text{update}}$ and $r^{\text{exit}}$ within end-to-end RL, rewarding the correct updating and exiting behaviors respectively. Experiments on various long-context reasoning tasks demonstrate the effectiveness and efficiency of GRU-Mem, which generally outperforms the vanilla MemAgent with up to 400\% times inference speed acceleration.

cs.CL↗

Advanced Multimodal Learning for Seizure Detection and Prediction: Concept, Challenges, and Future Directions

Epilepsy is a chronic neurological disorder characterized by recurrent unprovoked seizures, affects over 50 million people worldwide, and poses significant risks, including sudden unexpected death in epilepsy (SUDEP). Conventional unimodal approaches, primarily reliant on electroencephalography (EEG), face several key challenges, including low SNR, nonstationarity, inter- and intrapatient heterogeneity, portability, and real-time applicability in clinical settings. To address these issues, a comprehensive survey highlights the concept of advanced multimodal learning for epileptic seizure detection and prediction (AMLSDP). The survey presents the evolution of epileptic seizure detection (ESD) and prediction (ESP) technologies across different eras. The survey also explores the core challenges of multimodal and non-EEG-based ESD and ESP. To overcome the key challenges of the multimodal system, the survey introduces the advanced processing strategies for efficient AMLSDP. Furthermore, this survey highlights future directions for researchers and practitioners. We believe this work will advance neurotechnology toward wearable and imaging-based solutions for epilepsy monitoring, serving as a valuable resource for future innovations in this domain.

cs.NE↗

Analysis of Collaboration in CS Prizewinning with a Nobel-Turing Comparison

In the scientific community, prizes play a pivotal role in shaping research trajectories by conferring credibility and offering financial incentives to researchers. Yet, we know little about the relationship between academic collaborations and prizewinning. By analyzing over 100 scientific prizes and the collaboration behaviors of over 5,000 prizewinners in CS, we find that prizewinners collaborate earlier and more frequently with other prizewinners than researchers who have not yet received similar recognition. Moreover, CS researchers across age groups collaborate more with prizewinners after winning their first prize, and collaborating with prizewinners after their first win increases the likelihood of the collaborator winning an award. We find that recipients of general CS prizes collaborate more than recipients of more specialized prizes, who collaborate less frequently. With Coarsened Exact Matching (CEM) and regression, we find an increase in prizewinning odds with strength of prizewinner collaboration. We examine the context of recent Nobel Prizes going to CS researchers by showing how an increasing share of Physics awards go to Physics-CS collaborations, and contrast Nobel-Turing winning author's trajectories. Our findings shed light on the relationship between prizewinning and collaboration.

cs.SI↗

Virtual Width Networks

We introduce Virtual Width Networks (VWN), a framework that delivers the benefits of wider representations without incurring the quadratic cost of increasing the hidden size. VWN decouples representational width from backbone width, expanding the embedding space while keeping backbone compute nearly constant. In our large-scale experiment, an 8-times expansion accelerates optimization by over 2 times for next-token and 3 times for next-2-token prediction. The advantage amplifies over training as both the loss gap grows and the convergence-speedup ratio increases, showing that VWN is not only token-efficient but also increasingly effective with scale. Moreover, we identify an approximately log-linear scaling relation between virtual width and loss reduction, offering an initial empirical basis and motivation for exploring virtual-width scaling as a new dimension of large-model efficiency.

cs.LG↗

Structure-Attribute Transformations with Markov Chain Boost Graph Domain Adaptation

Graph domain adaptation has gained significant attention in label-scarce scenarios across different graph domains. Traditional approaches to graph domain adaptation primarily focus on transforming node attributes over raw graph structures and aligning the distributions of the transformed node features across networks. However, these methods often struggle with the underlying structural heterogeneity between distinct graph domains, which leads to suboptimal distribution alignment. To address this limitation, we propose Structure-Attribute Transformation with Markov Chain (SATMC), a novel framework that sequentially aligns distributions across networks via both graph structure and attribute transformations. To mitigate the negative influence of domain-private information and further enhance the model's generalization, SATMC introduces a private domain information reduction mechanism and an empirical Wasserstein distance. Theoretical proofs suggest that SATMC can achieve a tighter error bound for cross-network node classification compared to existing graph domain adaptation methods. Extensive experiments on nine pairs of publicly available cross-domain datasets show that SATMC outperforms state-of-the-art methods in the cross-network node classification task. The code is available at https://github.com/GiantZhangYT/SATMC.

cs.LG↗

Large Language Models Meet Graph Neural Networks: A Perspective of Graph Mining

Graph mining is an important area in data mining and machine learning that involves extracting valuable information from graph-structured data. In recent years, significant progress has been made in this field through the development of graph neural networks (GNNs). However, GNNs are still deficient in generalizing to diverse graph data. Aiming to this issue, Large Language Models (LLMs) could provide new solutions for graph mining tasks with their superior semantic understanding. In this review, we systematically review the combination and application techniques of LLMs and GNNs and present a novel taxonomy for research in this interdisciplinary field, which involves three main categories: GNN-driving-LLM, LLM-driving-GNN, and GNN-LLM-co-driving. Within this framework, we reveal the capabilities of LLMs in enhancing graph feature extraction as well as improving the effectiveness of downstream tasks such as node classification, link prediction, and community detection. Although LLMs have demonstrated their great potential in handling graph-structured data, their high computational requirements and complexity remain challenges. Future research needs to continue to explore how to efficiently fuse LLMs and GNNs to achieve more powerful graph learning and reasoning capabilities and provide new impetus for the development of graph mining techniques.

cs.LG↗

Deterministic vortices evolving from partially coherent fields

It has long been assumed that there is an intrinsic conflict between optical vortices and partial coherence, in that deterministic phase vortices do not appear in partially coherent fields. We demonstrate, however, that it is possible to construct a beam that has no deterministic vortices in the source plane yet evolves a deterministic vortex at a specified propagation distance.

physics.optics↗

An Adaptive Plug-and-Play Network for Few-Shot Learning

Few-shot learning (FSL) requires a model to classify new samples after learning from only a few samples. While remarkable results are achieved in existing methods, the performance of embedding and metrics determines the upper limit of classification accuracy in FSL. The bottleneck is that deep networks and complex metrics tend to induce overfitting in FSL, making it difficult to further improve the performance. Towards this, we propose plug-and-play model-adaptive resizer (MAR) and adaptive similarity metric (ASM) without any other losses. MAR retains high-resolution details to alleviate the overfitting problem caused by data scarcity, and ASM decouples the relationship between different metrics and then fuses them into an advanced one. Extensive experiments show that the proposed method could boost existing methods on two standard dataset and a fine-grained datasets, and achieve state-of-the-art results on mini-ImageNet and tiered-ImageNet.

cs.CV↗

Unsupervised Time-Aware Sampling Network with Deep Reinforcement Learning for EEG-Based Emotion Recognition

Recognizing human emotions from complex, multivariate, and non-stationary electroencephalography (EEG) time series is essential in affective brain-computer interface. However, because continuous labeling of ever-changing emotional states is not feasible in practice, existing methods can only assign a fixed label to all EEG timepoints in a continuous emotion-evoking trial, which overlooks the highly dynamic emotional states and highly non-stationary EEG signals. To solve the problems of high reliance on fixed labels and ignorance of time-changing information, in this paper we propose a time-aware sampling network (TAS-Net) using deep reinforcement learning (DRL) for unsupervised emotion recognition, which is able to detect key emotion fragments and disregard irrelevant and misleading parts. Extensive experiments are conducted on three public datasets (SEED, DEAP, and MAHNOB-HCI) for emotion recognition using leave-one-subject-out cross-validation, and the results demonstrate the superiority of the proposed method against previous unsupervised emotion recognition methods.

cs.HC↗

Universal Urban Spreading Pattern of COVID-19 and Its Underlying Mechanism

Currently, the global situation of COVID-19 is aggravating, pressingly calling for efficient control and prevention measures. Understanding spreading pattern of COVID-19 has been widely recognized as a vital step for implementing non-pharmaceutical measures. Previous studies investigated such an issue in large-scale (e.g., inter-country or inter-state) scenarios while urban spreading pattern still remains an open issue. Here, we fill this gap by leveraging the trajectory data of 197,808 smartphone users (including 17,808 anonymous confirmed cases) in 9 cities in China. We find a universal spreading pattern in all cities: the spatial distribution of confirmed cases follows a power-law-like model and the spreading centroid is time-invariant. Moreover, we reveal that human mobility in a city drives the spatialtemporal spreading process: long average travelling distance results in a high growth rate of spreading radius and wide spatial diffusion of confirmed cases. With such insight, we adopt Kendall model to simulate urban spreading of COVID-19 that can well fit the real spreading process. Our results unveil the underlying mechanism behind the spatial-temporal urban evolution of COVID-19, and can be used to evaluate the performance of mobility restriction policies implemented by many governments and to estimate the evolving spreading situation of COVID-19.

physics.soc-ph↗

Correlation-induced orbital angular momentum changes

We demonstrate that the orbital angular momentum flux density of a paraxial light beam can change on propagation in free space. These changes are entirely due to the spatial coherence state of the source, and the effect is analogous to correlation-induced changes in the intensity, spectrum and polarization of a beam. The use of the source coherence state to control the width, shape, and transverse shifts of the OAM flux density is demonstrated with numerical examples.

physics.optics↗

A Tensor-Based Framework for Studying Eigenvector Multicentrality in Multilayer Networks

Centrality is widely recognized as one of the most critical measures to provide insight in the structure and function of complex networks. While various centrality measures have been proposed for single-layer networks, a general framework for studying centrality in multilayer networks (i.e., multicentrality) is still lacking. In this study, a tensor-based framework is introduced to study eigenvector multicentrality, which enables the quantification of the impact of interlayer influence on multicentrality, providing a systematic way to describe how multicentrality propagates across different layers. This framework can leverage prior knowledge about the interplay among layers to better characterize multicentrality for varying scenarios. Two interesting cases are presented to illustrate how to model multilayer influence by choosing appropriate functions of interlayer influence and design algorithms to calculate eigenvector multicentrality. This framework is applied to analyze several empirical multilayer networks, and the results corroborate that it can quantify the influence among layers and multicentrality of nodes effectively.

physics.soc-ph↗