SearcharxivSearch

arXiv subjects

Zexing Zhao

Publications and source records attributed to Zexing Zhao.

7 recordsLinked to original sources

Towards Automatic Evolution Tree Generation from Citation Graphs

Surveys remain the primary way researchers grasp the lineage of methods within an AI subfield, but they scale poorly against the current rate of publication. Existing taxonomy-induction methods are largely leaf-bound and time-agnostic; they tend to force transitional papers into mature leaves and can create topological inversions between ancestors and descendants. We propose EvoTree, a staged framework that decouples conceptual backbone learning from temporal refinement: a graph-aware encoder with distribution-based hierarchical clustering yields a stable taxonomy backbone; temporal fine-tuning then re-attaches marginal papers to internal nodes under monotonic-path constraints; a final LLM pass labels concepts without altering the topology. We release the first annotated benchmark for this task across 11 AI subfields. EvoTree attains the highest NMI and citation-direction accuracy among all baselines and the best concept purity on the annotated benchmark, and is the only method with non-trivial marginal-paper detection on the annotated set.

cs.CL

Graph Navier Stokes Networks

Graph Neural Networks (GNNs) have emerged as a cornerstone of deep learning, with most existing methods rooted in graph signal processing and diffusion equations to model message passing. However, these approaches inherently suffer from the oversmoothing problem, where node features become indistinguishable as the network depth increases. Inspired by the Navier Stokes equations, we introduce Graph Navier Stokes Networks (GNSN), a novel architecture that transcends conventional diffusion-based message passing by incorporating convection into graph structures. GNSN defines a dynamic velocity field on the graph to govern convection, enabling more efficient and direct message propagation. By adaptively balancing convection and diffusion, GNSN is able to efficiently handle datasets with varying levels of homophily. Extensive evaluations across twelve real-world datasets demonstrate that GNSN consistently outperforms state-of-the-art baselines in classification accuracy. Moreover, experimental results further emphasize its effectiveness in alleviating the oversmoothing problem.

cs.LG

Optically locked low-noise photonic microwave oscillator

The next-generation sensing and communication applications rely on high-frequency microwave generation with low-noise. The microwave photonic technology is promising by the practical application is limited by its complex architecture so far. Here, we demonstrate an optically locked low-noise photonic microwave oscillator, so that all the optical components are packaged within a small module of 166 mL, and low noise microwave generation is achieved at 10.4 GHz with single-sideband phase noise of -54 dBc/Hz at 10 Hz, -141 dBc/Hz at 10 kHz, and -162 dBc/Hz at 10 MHz offset. Above performance arises from a dual-laser self-injection-locking scheme to a single Fabry-Perot cavity with high Q exceeding 10^8, with over 20 dB common-mode noise suppression. The low-noise nature of such reference is coherently transferred to the X-band through a high-performance TFLN electro-optic comb chip, thereby overcoming long-standing barriers in photonic microwave integration to enable truly field-deployable low-noise microwave generation.

physics.optics

HOP: Heterogeneous Topology-based Multimodal Entanglement for Co-Speech Gesture Generation

Co-speech gestures are crucial non-verbal cues that enhance speech clarity and expressiveness in human communication, which have attracted increasing attention in multimodal research. While the existing methods have made strides in gesture accuracy, challenges remain in generating diverse and coherent gestures, as most approaches assume independence among multimodal inputs and lack explicit modeling of their interactions. In this work, we propose a novel multimodal learning method named HOP for co-speech gesture generation that captures the heterogeneous entanglement between gesture motion, audio rhythm, and text semantics, enabling the generation of coordinated gestures. By leveraging spatiotemporal graph modeling, we achieve the alignment of audio and action. Moreover, to enhance modality coherence, we build the audio-text semantic representation based on a reprogramming module, which is beneficial for cross-modality adaptation. Our approach enables the trimodal system to learn each other's features and represent them in the form of topological entanglement. Extensive experiments demonstrate that HOP achieves state-of-the-art performance, offering more natural and expressive co-speech gesture generation. More information, codes, and demos are available here: https://star-uu-wang.github.io/HOP/

cs.CV

High-gain optical parametric amplification with a continuous-wave pump using a domain-engineered thin-film lithium niobate waveguide

While thin film lithium niobate (TFLN) is known for efficient signal generation, on-chip signal amplification remains challenging from fully integrated optical communication circuits. Here we demonstrate the continuous-wave-pump optical parametric amplification (OPA) using an x-cut domain-engineered TFLN waveguide, with high gain over the telecom band up to 13.9 dB, and test it for high signal-to-noise ratio signal amplification using a commercial optical communication module pair. Fabricated in wafer scale using common process as devices including modulators, this OPA device marks an important step in TFLN photonic integration.

physics.optics

Contrastive Dual-Interaction Graph Neural Network for Molecular Property Prediction

Molecular property prediction is a key component of AI-driven drug discovery and molecular characterization learning. Despite recent advances, existing methods still face challenges such as limited ability to generalize, and inadequate representation of learning from unlabeled data, especially for tasks specific to molecular structures. To address these limitations, we introduce DIG-Mol, a novel self-supervised graph neural network framework for molecular property prediction. This architecture leverages the power of contrast learning with dual interaction mechanisms and unique molecular graph enhancement strategies. DIG-Mol integrates a momentum distillation network with two interconnected networks to efficiently improve molecular characterization. The framework's ability to extract key information about molecular structure and higher-order semantics is supported by minimizing loss of contrast. We have established DIG-Mol's state-of-the-art performance through extensive experimental evaluation in a variety of molecular property prediction tasks. In addition to demonstrating superior transferability in a small number of learning scenarios, our visualizations highlight DIG-Mol's enhanced interpretability and representation capabilities. These findings confirm the effectiveness of our approach in overcoming challenges faced by traditional methods and mark a significant advance in molecular property prediction.

cs.LG

Passively stable 0.7-octave microcombs in thin-film lithium niobate microresonators

Optical frequency comb based on microresonator (microcomb) is an integrated coherent light source and has the potential to promise a high-precision frequency standard, and self-reference and long-term stable microcomb is the key to this realization. Here, we demonstrated a 0.7-octave spectrum Kerr comb via dispersion engineering in a thin film lithium niobate microresonator, and the single soliton state can be accessed passively with long-term stability over 3 hours. With such a robust broadband coherent comb source using thin film lithium niobate, fully stabilized microcomb can be expected for massive practical applications.

physics.optics