SearcharxivSearch

arXiv subjects

Shuang Wang

Publications and source records attributed to Shuang Wang.

At least 19 recordsLinked to original sources

DKDNet: Dual Knowledge and Data-Driven Network for Cross-Domain Automatic Modulation Classification

The dynamics of communication environments induce significant distribution shifts across domains, challenging the generalization of deep learning-based automatic modulation classification (AMC) models. While existing UDA methods alleviate this problem by aligning source and target features, they give limited consideration to modulation-specific structures that remain informative across domain conditions. In this paper, we consider signal prior knowledge, grounded in communication protocols and physical principles, as a potential way to enhance cross-domain representation learning. Given that different priors may vary in modulation discriminability, domain stability, and complementarity, this paper first analyzes five commonly adopted signal representations that instantiate different signal priors. From them, in-phase/quadrature (IQ), amplitude--phase (AP), and autocorrelation function (ACF) are selected as compact prior-guided inputs. Based on that, a dual knowledge and data-driven network (DKDNet) is proposed for cross-domain AMC. The multi-representation feature encoder (MRFE) and dynamic lightweight fusion unit (DLFU) are designed to achieve unified representation learning and adaptive feature fusion, and the resulting fused features are optimized with modulation classification and adversarial domain alignment objectives. Experiments on both simulated and public datasets validate the rationality of the prior selection and demonstrate the superiority of the proposed method.

eess.SP

FedMPT: Federated Multi-label Prompt Tuning of Vision-Language Models

Multi-Label Recognition (MLR) based on Vision-Language Models (VLMs) aims to leverage their pre-trained knowledge to better adapt complex recognition scenarios, thereby enhancing model robustness. However, for realistic decentralized applications requiring federated learning, adapting VLMs to each client that possesses private and heterogeneous data can cause the model to overfit spurious label correlations, consequently triggering irrelevant categories when encountering new samples. To tackle this problem, we reconsider the federated learning for MLR with a causal model, in which we adopt a front-door adjustment and decouple the MLR modeling process by intermediate variables that magnify the oracle label co-occurrence. Guided by our analysis, we propose our FedMPT, the first method specifically designed for federated MLR. The core idea of FedMPT is to leverage generalizable conditions to steer federated MLR to mitigate erroneous label activations. To achieve this, FedMPT introduces an Large Language Model (LLM)-driven pipeline to decipher the underlying conditions that govern label dependencies. Furthermore, we introduce an optimal transport between the condition-enriched prompts and the image patches to uncover multiple region-level semantics. Finally, we generate synergistic predictions from different conditions with a crafted gating mechanism. Experiments on multiple benchmark datasets show that our proposed approach achieves competitive results and outperforms SOTA methods under varied settings.

cs.AI

Modulation Consistency-based Contrastive Learning for Self-Supervised Automatic Modulation Classification

Deep learning-based AMC methods have achieved remarkable performance, but their practical deployment remains constrained by the high cost of labeled data. Although self-supervised learning (SSL) reduces the reliance on labels, existing SSL-based AMC methods often rely on task-agnostic pretext objectives misaligned with modulation classification, leading to representations entangled with nuisance factors such as symbol, channel, and noise. In this paper, we identify intra-instance modulation consistency as a task-aware structural prior, whereby different temporal segments of the same signal may differ in waveform while preserving the same modulation type, thus providing a principled cue for task-aligned self-supervision. Based on this prior, we propose Mod-CL, a Modulation consistency-based Contrastive Learning framework that constructs positive pairs from different temporal segments of the same signal instance, to encourage the model to learn shared modulation information while suppressing nuisance variations. We further develop a contrastive objective tailored to Mod-CL, which jointly exploits temporal segmentation and data augmentation to pull together views sharing the same modulation semantics while avoiding supervisory conflicts within each signal instance. Extensive experiments on RadioML datasets show that Mod-CL consistently outperforms strong baselines, especially in low-label regimes, achieving substantial improvements in linear probing accuracy.

eess.SP

TAR: Text Semantic Assisted Cross-modal Image Registration Framework for Optical and SAR Images

Existing deep learning-based methods can capture shared features from optical and synthetic aperture radar (SAR) images for spatial alignment. However, optical-SAR registration remains challenging under large geometric deformations, because the model needs to simultaneously handle cross-modal appearance discrepancies and complex spatial transformations. To address this issue, this paper proposes a text semantic-assisted cross-modal image registration framework, named TAR, for optical and SAR images. TAR exploits text semantic priors from remote sensing scenes and land-cover categories to alleviate the modality gap and enhance cross-modal feature learning. TAR consists of three components: a multi-scale visual feature learning (MSFL) module, a text-assisted feature enhancement (TAFE) module, and a coarse-to-fine dense matching (CFDM) module. MSFL extracts multi-scale visual features from optical and SAR images. TAFE constructs text descriptors related to remote sensing scenes and land-cover objects, and uses a frozen RemoteCLIP text encoder to extract text features. These text features are introduced through visual-text interaction to enhance high-level visual features for more reliable coarse matching. CFDM then establishes coarse correspondences based on the enhanced high-level features and refines the matched locations using low-level features. Experimental results on cross-modal remote sensing images demonstrate the effectiveness of TAR, which achieves stronger matching performance than several state-of-the-art methods and yields significant gains under large geometric deformations.

cs.CV

BGG: Bridging the Geometric Gap between Cross-View images by Vision Foundation Model Adaptation for Geo-Localization

Geometric differences between cross-view images, such as drone and satellite views, significantly increase the challenge of Cross-View Geo-Localization (CVGL), which aims to acquire the geolocation of images by image retrieval. To further enhance the CVGL performance, this paper proposes a parameter-efficient adaptation framework for bridging the geometric gap across images based on the vision foundation model (VFM) (e.g., DINOv3), termed BGG. BGG not only effectively leverages the general visual representations of VFM and captures the robust and consistent features from cross-view images, but also utilizes the generalization capabilities of the VFM, significantly improving the CVGL performance. It mainly contains a Multi-granularity Feature Enhancement Adapter (MFEA) and a Frequency-Aware Structural Aggregation (FASA) module. Specifically, MFEA enhances the scale adaptability and viewpoint robustness of features by multi-level dilated convolutions, effectively bridging the cross-view geometric gap with small training costs. Additionally, considering the [CLS] token lacks spatial details for precise image retrieval and localization, the FASA module modulates patch tokens in the frequency domain and performs adaptive aggregation for local structural feature enhancement. Finally, BGG fuses the enhanced local features with the [CLS] token for more accurate CVGL. Extensive experiments on University-1652 and SUES-200 datasets demonstrate that BGG has significant advantages over other methods and achieves state-of-the-art localization performance with low training costs.

cs.CV

Learning What Matters Now: Dynamic Preference Inference under Contextual Shifts

Humans often juggle multiple, sometimes conflicting objectives and shift their priorities as circumstances change, rather than following a fixed objective function. In contrast, most computational decision-making and multi-objective RL methods assume static preference weights or a known scalar reward. In this work, we study sequential decision-making problem when these preference weights are unobserved latent variables that drift with context. Specifically, we propose Dynamic Preference Inference (DPI), a cognitively inspired framework in which an agent maintains a probabilistic belief over preference weights, updates this belief from recent interaction, and conditions its policy on inferred preferences. We instantiate DPI as a variational preference inference module trained jointly with a preference-conditioned actor-critic, using vector-valued returns as evidence about latent trade-offs. In queueing, maze, and multi-objective continuous-control environments with event-driven changes in objectives, DPI adapts its inferred preferences to new regimes and achieves higher post-shift performance than fixed-weight and heuristic envelope baselines.

cs.AI

A Tug-of-War Between Baroclinic Eddies and Convection: Implications for Icy Moon Oceans

In many geophysical and planetary environments, such as Earth's ocean and atmosphere as well as subsurface oceans of icy satellites, convection driven by bottom geothermal heating usually coexists with baroclinic eddies driven by lateral buoyancy/temperature gradients. These processes compete against each other, with convection destabilizing the stratification and baroclinic eddies re-stabilizing it, thereby controlling whether the bottom heat flux is significantly redistributed as it is transmitted to the upper surface. Using scaling analysis and numerical simulations, we show that a stratified layer persists near the upper surface up to ${\rm Ra}_{v}\sim {\rm Ra}_h^{5/2}$, where ${\rm Ra}_h\equiv \Delta b_0/(L_zf^2)$ measures the imposed upper-surface buoyancy contrast $\Delta b_0$ and ${\rm Ra}_v\equiv B_0/(L_z^2f^3)$ measures the strength of the bottom buoyancy flux $B_0$, $L_z$ is the domain depth and $f$ is the Coriolis parameter. For ${\rm Ra}_v<{\rm Ra}_h^{5/2}$, baroclinic eddies dominate over convection, maintain the upper stratified layer, and completely deflect the bottom buoyancy/heat input into meridional transport. In contrast, when ${\rm Ra}_v>{\rm Ra}_h^{5/2}$, convective plumes penetrate the stratification and transport buoyancy/heat vertically with negligible deflection. Building on these results, we further propose a scaling law for the meridional buoyancy/heat transport in this system. Applications to icy satellites are discussed.

physics.ao-ph

Generalizable Knowledge Distillation from Vision Foundation Models for Semantic Segmentation

Knowledge distillation (KD) has been widely applied in semantic segmentation to compress large models, but conventional approaches primarily preserve in-domain accuracy while neglecting out-of-domain generalization, which is essential under distribution shifts. This limitation becomes more severe with the emergence of vision foundation models (VFMs): although VFMs exhibit strong robustness on unseen data, distilling them with conventional KD often compromises this ability. We propose Generalizable Knowledge Distillation (GKD), a multi-stage framework that explicitly enhances generalization. GKD decouples representation learning from task learning. In the first stage, the student acquires domain-agnostic representations through selective feature distillation, and in the second stage, these representations are frozen for task adaptation, thereby mitigating overfitting to visible domains. To further support transfer, we introduce a query-based soft distillation mechanism, where student features act as queries to teacher representations to selectively retrieve transferable spatial knowledge from VFMs. Extensive experiments on five domain generalization benchmarks demonstrate that GKD consistently outperforms existing KD methods, achieving average gains of +1.9% in foundation-to-foundation (F2F) and +10.6% in foundation-to-local (F2L) distillation. The code will be available at https://github.com/Younger-hua/GKD.

cs.CV

SpecFuse: A Spectral-Temporal Fusion Predictive Control Framework for UAV Landing on Oscillating Marine Platforms

Autonomous landing of Uncrewed Aerial Vehicles (UAVs) on oscillating marine platforms is severely constrained by wave-induced multi-frequency oscillations, wind disturbances, and prediction phase lags in motion prediction. Existing methods either treat platform motion as a general random process or lack explicit modeling of wave spectral characteristics, leading to suboptimal performance under dynamic sea conditions. To address these limitations, we propose SpecFuse: a novel spectral-temporal fusion predictive control framework that integrates frequency-domain wave decomposition with time-domain recursive state estimation for high-precision 6-DoF motion forecasting of Uncrewed Surface Vehicles (USVs). The framework explicitly models dominant wave harmonics to mitigate phase lags, refining predictions in real time via IMU data without relying on complex calibration. Additionally, we design a hierarchical control architecture featuring a sampling-based HPO-RRT* algorithm for dynamic trajectory planning under non-convex constraints and a learning-augmented predictive controller that fuses data-driven disturbance compensation with optimization-based execution. Extensive validations (2,000 simulations + 8 lake experiments) show our approach achieves a 3.2 cm prediction error, 4.46 cm landing deviation, 98.7% / 87.5% success rates (simulation / real-world), and 82 ms latency on embedded hardware, outperforming state-of-the-art methods by 44%-48% in accuracy. Its robustness to wave-wind coupling disturbances supports critical maritime missions such as search and rescue and environmental monitoring. All code, experimental configurations, and datasets will be released as open-source to facilitate reproducibility.

cs.RO

Relativistic Position Verification with Coherent States

Determining the position of an entity is a fundamental prerequisite for nearly all activities. Classical means, however, have been proven incapable of providing secure position verification, meaning that a prover can mislead verifiers about its actual position. In this work, we propose and experimentally realize a secure position-verification protocol that leverages quantum optics and relativity within an information-theoretic framework. Using phase-randomized weak coherent states, two verifiers separated by 2 km securely verify the prover's position with an accuracy better than 75 meters. These results establish secure position-based authentication as a practical possibility, paving the way for applications in financial transactions, disaster response, and authenticated secure communications.

quant-ph

5-GHz chip-based quantum key distribution with 1Mbps secure key rate over 150 km

Quantum key distribution (QKD) enables secure communication by harnessing the fundamental principles of quantum physics, which inherently guarantee information-theoretic security and intrinsic resistance to quantum computing attacks. However, the secure key rate of QKD typically decreases exponentially with increasing channel distance. In this work, by developing a novel polarization-state preparation method, an ultra-low time-jitter laser source and superconducting nanowire single-photon detectors, we demonstrate a 5-GHz integrated QKD system featuring ultra-low quantum bit error rates (QBERs). The system achieves secure key rates of 1.076 Mbps at 150 km and 105 kbps at 200 km over standard single-mode fiber channels, respectively. Our system substantially enhances the secure key rate, enabling high-resolution video calls with one-time-pad encryption over intercity backbone QKD links. This work represents a significant step forward in the development of high-performance practical QKD systems.

quant-ph

Microcomb-driven large-scale fully connected quantum network

Fully connected quantum networks enable simultaneously connecting every user to every other user and are the most versatile and robust networking architecture. However, the scalability of such networks remains great challenge for practical applications. Here we construct a large-scale fully connected quantum network founded on two-photon Hong-Ou-Mandel (HOM) interference, where user-to-user security is guaranteed even with untrusted network provider. Using integrated soliton microcomb (SMC) and photonic encoding chips, we realize precise massive parallel frequency generation and locking, high-visibility HOM interferences and measurement-device-independent (MDI) quantum key distribution. The proposed architecture enables a 200-user fully connected quantum network over 200 kilometers with strict information-theoretic security via untrusted network provider. The implemented networking architecture paves the way for realizing large-scale fully connected MDI quantum networks across metropolitan and intercity regions.

quant-ph

CLNet: Cross-View Correspondence Makes a Stronger Geo-Localizationer

Image retrieval-based cross-view geo-localization (IRCVGL) aims to match images captured from significantly different viewpoints, such as satellite and street-level images. Existing methods predominantly rely on learning robust global representations or implicit feature alignment, which often fail to model explicit spatial correspondences crucial for accurate localization. In this work, we propose a novel correspondence-aware feature refinement framework, termed CLNet, that explicitly bridges the semantic and geometric gaps between different views. CLNet decomposes the view alignment process into three learnable and complementary modules: a Neural Correspondence Map (NCM) that spatially aligns cross-view features via latent correspondence fields; a Nonlinear Embedding Converter (NEC) that remaps features across perspectives using an MLP-based transformation; and a Global Feature Recalibration (GFR) module that reweights informative feature channels guided by learned spatial cues. The proposed CLNet can jointly capture both high-level semantics and fine-grained alignments. Extensive experiments on four public benchmarks, CVUSA, CVACT, VIGOR, and University-1652, demonstrate that our proposed CLNet achieves state-of-the-art performance while offering better interpretability and generalizability.

cs.CV

Photorefractive-based on-chip optical power limiter against light-injection attacks in quantum key distribution

Light-injection attacks pose critical security threats to quantum key distribution (QKD) systems. Conventional countermeasures, such as isolators, filters, and optical power monitoring, suffer from limited on-chip compatibility and inherent security vulnerabilities. To overcome these limitations, we propose and experimentally demonstrate an integrated attack sensing and automatic response unit utilizing the photorefractive effect in a thin-film lithium niobate microring resonator. The unit provides a rejection ratio exceeding 25 dB against non-resonant injected light. Under resonant attacks with power levels above tens of microwatts, the unit autonomously attenuates the signal transmission, with 14 dB attenuation measured at the maximum tested attack power of 10 dBm, leading to a significant suppression of the secure key rate. We further verify its response to pulsed light injection and incorporate possible residual leakage associated with finite response time into the key-rate analysis. This work provides a highly sensitive, broadband, and fully on-chip defense mechanism that significantly enhances the physical-layer security of QKD systems against light-injection attacks.

quant-ph

Revisiting the Hubble tension problem in the framework of holographic dark energy

The Hubble tension problem is one of the most significant challenges in modern cosmology. In this paper, we study the Hubble tension problem in the framework of holographic dark energy (HDE). To perform a systematic and comprehensive analysis, we select six representative theoretical models from all four categories of HDE. For the observational data, we adopt the Baryon Acoustic Oscillation (BAO) data from the Dark Energy Spectroscopic Instrument (DESI) Data Release 2 (DR2) along with a collection of alternative BAO measurements, Cosmic Microwave Background (CMB) distance priors from $Planck$ 2018, and type Ia supernovae (SN) data from the PantheonPlus, Union3, and DESY5 compilations. We find that HDE models that employ the Hubble scale or its combinations as the infrared (IR) cutoff cannot alleviate the Hubble tension problem. In contrast, HDE models that employ the future event horizon as the IR cutoff can partially mitigate the Hubble tension problem. It must be stressed that these two key conclusions hold true for cases of adopting different theoretical HDE models and different observational data. Our findings advocate for further exploration of HDE models using other types of cosmological observations.

astro-ph.CO

Human-AI Co-Embodied Intelligence for Scientific Experimentation and Manufacturing

Scientific experimentation and manufacturing rely on prolonged protocol development and complex, multi-step implementation, which require continuous human expertise for precise execution and decision-making, limiting interpretability and scalability. Here, we introduce human-artificial intelligence (AI) co-embodied intelligence, a new form of physical AI that unites human researchers, agentic AI, and wearable hardware. In this paradigm, humans provide precise execution, while agentic AI contributes contextual reasoning, adaptive planning, and analysis. The wearable interface continuously captures experimentation and manufacturing, facilitating seamless communication between humans and AI. We instantiate this paradigm in a microfabrication cleanroom, leading to the agentic-physical experimentation (APEX) system which understands fabrication procedure with accuracy 51% higher than state-of-the-art multimodal large language models/vision language models (LLMs/VLMs), detects and corrects fabrication errors in real-time, and transfers procedural expertise to novice users. Critically, APEX system enables the co-development of fabrication protocols in cleanrooms, overcoming the incompatibility of elastomeric materials in standard microfabrication processes and enabling previously unattainable fabrication outcomes, as demonstrated by the wafer-scale realization of brain-level soft neural probe capable of single-unit-resolution neural recording. These results establish the human-AI co-embodied intelligence that extends agentic reasoning beyond computation into the physical domain, transforming scientific experimentation and manufacturing into autonomous, traceable, interpretable and scalable processes.

cs.AI

OpenEM: Large-scale multi-structural 3D datasets for electromagnetic methods

Electromagnetic (EM) methods, owing to their efficiency and non-invasive nature, have become one of the most widely used techniques in geological exploration. Nevertheless, data processing for these methods remains highly time-consuming and labor-intensive. With the remarkable success of deep learning, applying such techniques to EM methods has emerged as a promising research direction to overcome the limitations of conventional approaches. The effectiveness of deep learning methods depends heavily on the quality of datasets, which directly influences model performance and generalization ability. Existing application studies often construct datasets from random one-dimensional or structurally simple threedimensional (3D) models, which fail to represent the complexity of real geological environments. Furthermore, the absence of standardized, publicly available 3D geoelectric datasets continues to hinder progress in deep learning based EM exploration. To address these limitations, we present OpenEM, a large-scale, multi-structural 3D geoelectric dataset that encompasses a broad range of geologically plausible subsurface structures. OpenEM consists of nine categories of geoelectric models, spanning from simple configurations with anomalous bodies in half-space to more complex structures such as flat layers, folded layers, flat faults, curved faults and their corresponding variants with anomalous bodies. In addition, we provide a 3D model generator that enables fully controllable 3D model construction, allowing flexible and extensible augmentation of OpenEM. OpenEM provides a unified, comprehensive, and large-scale dataset for common EM exploration systems to accelerate the application of deep learning in electromagnetic methods. The complete dataset and 3D model generator is publicly available at https://doi.org/10.5281/zenodo.17141981.

cs.LG

Fast qubit-based frequency recovery algorithm for quantum key distribution

Clock synchronization serves as a foundational subsystem in quantum key distribution (QKD). The recently proposed Qubit-based synchronization (Qubit4Sync) has opportunities in eliminating additional cost, noise, and potential side channels. It offers a promising alternative to dedicated synchronization hardware. However, the current frequency recovery process in Qubit4Sync requires high data throughput and computational speed, limiting practical use. To overcome these issues, we developed a fast frequency recovery algorithm that increases the recovery rate by orders of magnitude and remains robust under bad signal-to-noise ratio (SNR). This enables Qubit4Sync to operate effectively in mainstream gated-mode QKD systems. We further establish a theoretical model for frequency recovery, showing that our algorithm is robust against disturbances like dead time, jitter, and afterpulse. A frequency-domain SNR calculation method is also provided to guide parameter design for specific experimental conditions. This work opens the door to practical Qubit4Sync deployment in general QKD systems.

quant-ph