SearcharxivSearch

arXiv subjects

Jian Xiao

Publications and source records attributed to Jian Xiao.

At least 19 recordsLinked to original sources

Shift-MoE-Based DJSCC for CSI Feedback in Multi-User Pinching-Antenna Systems

In frequency-division duplexing systems, the performance gains of pinching-antenna systems (PASS) critically depend on accurate channel state information (CSI) at the base station. However, PASS CSI exhibits structured correlations over the waveguide-antenna grid and pronounced heterogeneity across users, making conventional fixed feedback mappings difficult to generalize. To address this challenge, this letter proposes an end-to-end CSI feedback scheme over a noisy uplink feedback link based on deep joint source-channel coding, termed Shift-based Mixture-of-Experts (Shift-MoE). Specifically, Shift-MoE leverages channel-grouped one-step shift operations to capture grid dependencies without global attention, and employs a gated multilayer perceptron mixture-of-experts module to adapt to heterogeneous CSI statistics across users. Numerical results demonstrate that the proposed Shift-MoE consistently outperforms representative learning-based CSI feedback baselines in normalized mean squared error and remains effective under different system parameter settings.

eess.SP

Channel Estimation for Rydberg Atomic Quantum Receivers: Unrolled Phase Retrieval from Holographic Snapshots

A model-driven deep learning framework is proposed for channel estimation in Rydberg atomic quantum receivers (RAQRs) based on the measurement of holographic snapshots. Specifically, we develop a Transformer-based unrolling architecture, termed URformer, to solve the non-linear biased phase retrieval problem, which is derived by unrolling a stabilized variant of the expectation-maximization Gerchberg-Saxton (EM-GS) algorithm. Each layer of the proposed URformer incorporates three trainable modules: 1) a learnable filter network that replaces the fixed Bessel kernel in the classic EM-GS algorithm; 2) a trainable gating mechanism that adaptively combines classic updates to ensure training stability; and 3) an efficient channel Transformer module that learns to correct residual errors by capturing non-local channel dependencies. Numerical results demonstrate that the proposed URformer significantly outperforms classic iterative algorithms and conventional black-box neural networks with less pilot overhead.

cs.IT

Channel Estimation for Flexible Intelligent Metasurfaces: From Model-Based Approaches to Neural Operators

Flexible intelligent metasurfaces (FIMs) offer a new solution for wireless communications by introducing morphological degrees of freedom, dynamically morphing their three-dimensional shape to ensure multipath signals interfere constructively. However, realizing the desired performance gains in FIM systems is critically dependent on acquiring accurate channel state information across a continuous and high-dimensional deformation space. Therefore, this paper investigates this fundamental channel estimation problem for FIM assisted millimeter-wave communication systems. First, we develop model-based frameworks that structure the problem as either function approximation using interpolation and kernel methods or as a sparse signal recovery problem that leverages the inherent angular sparsity of millimeter-wave channels. To further advance the estimation capability beyond explicit assumptions in model-based channel estimation frameworks, we propose a deep learning-based framework using a Fourier neural operator (FNO). By parameterizing a global convolution operator in the Fourier domain, we design an efficient FNO architecture to learn the continuous operator that maps FIM shapes to channel responses with mesh-independent properties. Furthermore, we exploit a hierarchical FNO (H-FNO) architecture to efficiently capture the multi-scale features across a hierarchy of spatial resolutions. Numerical results demonstrate that the proposed H-FNO significantly outperforms the model-based benchmarks in estimation accuracy and pilot efficiency. In particular, the interpretability analysis show that the proposed H-FNO learns an anisotropic spatial filter adapted to the physical geometry of FIM and is capable of accurately reconstructing the non-linear channel response across the continuous deformation space.

cs.IT

Wireless AI Evolution: From Statistical Learners to Electromagnetic-Guided Foundation Models

While initial applications of artificial intelligence (AI) in wireless communications over the past decade have demonstrated considerable potential using specialized models for targeted communication tasks, the revolutionary demands of sixth-generation (6G) networks for holographic communications, ubiquitous sensing, and native intelligence are propelling a necessary evolution towards AI-native wireless networks. The arrival of large AI models paves the way for the next phase of Wireless AI, driven by wireless foundation models (WFMs). In particular, pre-training on universal electromagnetic (EM) principles equips WFMs with the essential adaptability for a multitude of demanding 6G applications. However, existing large AI models face critical limitations, including pre-training strategies disconnected from EM-compliant constraints leading to physically inconsistent predictions, a lack of embedded understanding of wave propagation physics, and the inaccessibility of massive labeled datasets for comprehensive EM-aware training. To address these challenges, this article presents an electromagnetic information theory-guided self-supervised pre-training (EIT-SPT) framework designed to systematically inject EM physics into WFMs. The EIT-SPT framework aims to infuse WFMs with intrinsic EM knowledge, thereby enhancing their physical consistency, generalization capabilities across varied EM landscapes, and overall data efficiency. Building upon the proposed EIT-SPT framework, this article first elaborates on diverse potential applications in 6G scenarios of WFMs, then validates the efficacy of the proposed framework through illustrative case studies, and finally summarizes critical open research challenges and future directions for WFMs.

cs.IT

Rydberg Atomic Quantum Receivers for Wireless Communications: Two-Color vs. Three-Color Excitation

An efficient three-color (3C) laser excitation-based Rydberg atomic quantum receiver (RAQR) architecture is investigated for wireless communications, utilizing a five-level (5L) electronic transition mechanism. Specifically, the conventional two-color (2C) RAQR with the four-level (4L) excitation faces three fundamental obstacles: 1) high cost and engineering challenges due to the reliance on unstable short-wavelength lasers; 2) a fundamental sensitivity limit in thermal atoms caused by residual Doppler broadening; and 3) the inability to detect low-frequency bands due to the energy-level constraint of two-photon resonance. To address these challenges, this paper analyzes a 3C5L-RAQR architecture with all-red/infrared lasers, which not only solves the engineering cost issues but also enables effective Doppler cancellation and low-frequency detection by exploiting the three-photon resonance. Bridging atomic physics and communication theory, an end-to-end equivalent baseband signal model is derived. Furthermore, the performance of different RAQR architectures is evaluated in terms of sensitivity, achievable rate and spectrum access range. Moreover, we provide an exact numerical solution for practical RAQRs by employing the Liouvillian superoperator formalism. Numerical results demonstrate that the exhibited 3C5L-RAQR achieves superior sensitivity compared to the conventional 2C4L-RAQR and a classical antenna-based radio frequency receiver for weak-signal detection. Finally, the inherent sensitivity-rate trade-off is revealed, showing that the 3C5L-RAQR is more suitable for deployment in power-limited communication scenarios demanding broad spectrum access.

cs.IT

GenHOI: Generalized Hand-Object Pose Estimation with Occlusion Awareness

Generalized 3D hand-object pose estimation from a single RGB image remains challenging due to the large variations in object appearances and interaction patterns, especially under heavy occlusion. We propose GenHOI, a framework for generalized hand-object pose estimation with occlusion awareness. GenHOI integrates hierarchical semantic knowledge with hand priors to enhance model generalization under challenging occlusion conditions. Specifically, we introduce a hierarchical semantic prompt that encodes object states, hand configurations, and interaction patterns via textual descriptions. This enables the model to learn abstract high-level representations of hand-object interactions for generalization to unseen objects and novel interactions while compensating for missing or ambiguous visual cues. To enable robust occlusion reasoning, we adopt a multi-modal masked modeling strategy over RGB images, predicted point clouds, and textual descriptions. Moreover, we leverage hand priors as stable spatial references to extract implicit interaction constraints. This allows reliable pose inference even under significant variations in object shapes and interaction patterns. Extensive experiments on the challenging DexYCB and HO3Dv2 benchmarks demonstrate that our method achieves state-of-the-art performance in hand-object pose estimation.

cs.CV

Efficient Cross-Architecture Knowledge Transfer for Large-Scale Online User Response Prediction

Deploying new architectures in large-scale user response prediction systems incurs high model switching costs due to expensive retraining on massive historical data and performance degradation under data retention constraints. Existing knowledge distillation methods struggle with architectural heterogeneity and the prohibitive cost of transferring large embedding tables. We propose CrossAdapt, a two-stage framework for efficient cross-architecture knowledge transfer. The offline stage enables rapid embedding transfer via dimension-adaptive projections without iterative training, combined with progressive network distillation and strategic sampling to reduce computational cost. The online stage introduces asymmetric co-distillation, where students update frequently while teachers update infrequently, together with a distribution-aware adaptation mechanism that dynamically balances historical knowledge preservation and fast adaptation to evolving data. Experiments on three public datasets show that CrossAdapt achieves 0.27-0.43% AUC improvements while reducing training time by 43-71%. Large-scale deployment on Tencent WeChat Channels (~10M daily samples) further demonstrates its effectiveness, significantly mitigating AUC degradation, LogLoss increase, and prediction bias compared to standard distillation baselines.

cs.AI

Practical Channel Estimation for Pinching-Antenna Systems: Serial vs. Parallel and Downlink vs. Uplink?

The practical channel estimation for pinching-antenna networks is investigated, in which an electromagnetic-compliant in-waveguide transmission model is exhibited, incorporating bidirectional power splitting, cumulative power leakage, and waveguide attenuation. Based on this model, the paper investigates two antenna activation protocols for channel estimation: a serial protocol based on one-by-one antenna activation and a parallel protocol utilizing a binary S-Matrix activation. The serial protocol is characterized by its superior numerical stability but a lack of array gain, whereas the parallel protocol theoretically offers array gain but suffers from severe performance degradation due to structural crosstalk from the non-orthogonal S-Matrix and ill-conditioning from cumulative leakage. Furthermore, the paper analyzes the fundamental commonalities and asymmetries between uplink and downlink channel estimation in pinching-antenna systems. Numerical results demonstrate that 1) in an ideal lossless model, the parallel protocol is superior to the serial protocol due to the array gain from simultaneous energy collection in uplink transmission; 2) in a practical model with physical losses, the serial protocol outperforms the parallel protocol, as the performance of the parallel protocol is degraded by the numerical instability from cumulative leakage, which outweighs the benefit of array gain; 3) For downlink channel estimation, the serial protocol is more suitable because its strategy of concentrating the entire power budget on one measurement, while the parallel protocol is more suitable for the uplink as it can make full use of array gain.

cs.IT

Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval

Recent progress in text-video retrieval has been largely driven by contrastive learning. However, existing methods often overlook the effect of the modality gap, which causes anchor representations to undergo in-place optimization (i.e., optimization tension) that limits their alignment capacity. Moreover, noisy hard negatives further distort the semantics of anchors. To address these issues, we propose GARE, a Gap-Aware Retrieval framework that introduces a learnable, pair-specific increment $Δ_{ij}$ between text $t_i$ and video $v_j$, redistributing gradients to relieve optimization tension and absorb noise. We derive $Δ_{ij}$ via a multivariate first-order Taylor expansion of the InfoNCE loss under a trust-region constraint, showing that it guides updates along locally consistent descent directions. A lightweight neural module conditioned on the semantic gap couples increments across batches for structure-aware correction. Furthermore, we regularize $Δ$ through a variational information bottleneck with relaxed compression, enhancing stability and semantic consistency. Experiments on four benchmarks demonstrate that GARE consistently improves alignment accuracy and robustness, validating the effectiveness of gap-aware tension mitigation. Code is available at https://github.com/musicman217/GARE-text-video-retrieval.

cs.CV

Frequency-Selective Modeling and Analysis for OFDM-Integrated Wideband Pinching-Antenna Systems

This letter investigates the integration of pinching-antenna systems (PASS) with orthogonal frequency division multiplexing (OFDM) to ensure their compatibility and to explore the frequency-selective behavior inherent to PASS. First, an end-to-end channel model for OFDM PASS is proposed based on electromagnetic-compliant modeling of waveguides and coupled-mode theory, which includes frequency-dependent waveguide attenuation, dispersion and antenna coupling effect. Furthermore, a critical dependence of the OFDM cyclic prefix (CP) overhead on the proximity of the operating frequency to the waveguide cutoff is revealed. Moreover, the phase misalignment effect across subcarriers in OFDM PASS is derived for an approximate pinching antenna location strategy based on path loss minimization, which reveals the phase misalignment is exacerbated for wider bandwidths and larger array size. Numerical results show that: 1) frequency-selective effects in OFDM PASS lead to substantial variations in subcarrier achievable rates, highlighting the necessity of operating above the waveguide cutoff frequency for effective communications; 2) waveguide dispersion mandates considerable CP overhead when operating near the cutoff frequency, severely impacting the spectral efficiency of OFDM PASS; and 3) the gentle linear waveguide attenuation in a practical PASS significantly more advantageous than the severe logarithmic path loss characteristic of fixed-location antennas.

cs.IT

Learning Function-to-Function Mappings: A Fourier Neural Operator for Next-Generation MIMO Systems

Next-generation multiple-input multiple-output (MIMO) systems, characterized by extremely large-scale arrays, holographic surfaces, three-dimensional architectures, and flexible antennas, are poised to deliver unprecedented data rates, spectral efficiency and stability. However, these advancements introduce significant challenges for physical layer signal processing, stemming from complex near-field propagation, continuous aperture modeling, sub-wavelength antenna coupling effects, and dynamic channel conditions. Conventional model-based and deep learning approaches often struggle with the immense computational complexity and model inaccuracies inherent in these new regimes. This article proposes a Fourier neural operator (FNO) as a powerful and promising tool to address these challenges. The FNO learns function-to-function mappings between infinite-dimensional function spaces, making them exceptionally well-suited for modeling complex physical systems governed by partial differential equations based on electromagnetic wave propagation. We first present the fundamental principles of FNO, demonstrating its mesh-free nature and function-to-function ability to efficiently capture global dependencies in the Fourier domain. Furthermore, we explore a range of applications of FNO in physical-layer signal processing for next-generation MIMO systems. Representative case studies on channel modeling and estimation for novel MIMO architectures demonstrate the superior performance of FNO compared to state-of-the-art methods. Finally, we discuss open challenges and outline future research directions, positioning FNO as a promising technology for enabling the enormous potential of next-generation MIMO systems.

cs.IT

Two-Timescale Learning for Pilot-Free ISAC Systems

A pilot-free integrated sensing and communication (ISAC) system is investigated, in which phase-modulated continuous wave (PMCW) and non-orthogonal multiple access (NOMA) waveforms are co-designed to achieve simultaneous target sensing and data transmission. To enhance effective data throughput (i.e., Goodput) in PMCW-NOMA ISAC systems, we propose a deep learning-based receiver architecture, termed two-timescale Transformer (T3former), which leverages a Transformer architecture to perform joint channel estimation and multi-user signal detection without the need for dedicated pilot signals. By treating the deterministic structure of the PMCW waveform as an implicit pilot, the proposed T3former eliminates the overhead associated with traditional pilot-based methods. The proposed T3former processes the received PMCW-NOMA signals on two distinct timescales, where a fine-grained attention mechanism captures local features across the fast-time dimension, while a coarse-grained mechanism aggregates global spatio-temporal dependencies of the slow-time dimension. Numerical results demonstrate that the proposed T3former significantly outperforms traditional successive interference cancellation (SIC) receivers, which avoids inherent error propagation in SIC. Specifically, the proposed T3former achieves a substantially lower bit error rate and a higher Goodput, approaching the theoretical maximum capacity of a pilot-free system.

cs.IT

When Vision-Language Model (VLM) Meets Beam Prediction: A Multimodal Contrastive Learning Framework

As the real propagation environment becomes in creasingly complex and dynamic, millimeter wave beam prediction faces huge challenges. However, the powerful cross modal representation capability of vision-language model (VLM) provides a promising approach. The traditional methods that rely on real-time channel state information (CSI) are computationally expensive and often fail to maintain accuracy in such environments. In this paper, we present a VLM-driven contrastive learning based multimodal beam prediction framework that integrates multimodal data via modality-specific encoders. To enforce cross-modal consistency, we adopt a contrastive pretraining strategy to align image and LiDAR features in the latent space. We use location information as text prompts and connect it to the text encoder to introduce language modality, which further improves cross-modal consistency. Experiments on the DeepSense-6G dataset show that our VLM backbone provides additional semantic grounding. Compared with existing methods, the overall distance-based accuracy score (DBA-Score) of 0.9016, corresponding to 1.46% average improvement.

eess.SP

Deep Learning Based Joint Channel Estimation and Positioning for Sparse XL-MIMO OFDM Systems

This paper investigates joint channel estimation and positioning in near-field sparse extra-large multiple-input multiple-output (XL-MIMO) orthogonal frequency division multiplexing (OFDM) systems. To achieve cooperative gains between channel estimation and positioning, we propose a deep learning-based two-stage framework comprising positioning and channel estimation. In the positioning stage, the user's coordinates are predicted and utilized in the channel estimation stage, thereby enhancing the accuracy of channel estimation. Within this framework, we propose a U-shaped Mamba architecture for channel estimation and positioning, termed as CP-Mamba. This network integrates the strengths of the Mamba model with the structural advantages of U-shaped convolutional networks, enabling effective capture of local spatial features and long-range temporal dependencies of the channel. Numerical simulation results demonstrate that the proposed two-stage approach with CP-Mamba architecture outperforms existing baseline methods. Moreover, sparse arrays (SA) exhibit significantly superior performance in both channel estimation and positioning accuracy compared to conventional compact arrays.

eess.SP

Occlusion-Aware 3D Hand-Object Pose Estimation with Masked AutoEncoders

Hand-object pose estimation from monocular RGB images remains a significant challenge mainly due to the severe occlusions inherent in hand-object interactions. Existing methods do not sufficiently explore global structural perception and reasoning, which limits their effectiveness in handling occluded hand-object interactions. To address this challenge, we propose an occlusion-aware hand-object pose estimation method based on masked autoencoders, termed as HOMAE. Specifically, we propose a target-focused masking strategy that imposes structured occlusion on regions of hand-object interaction, encouraging the model to learn context-aware features and reason about the occluded structures. We further integrate multi-scale features extracted from the decoder to predict a signed distance field (SDF), capturing both global context and fine-grained geometry. To enhance geometric perception, we combine the implicit SDF with an explicit point cloud derived from the SDF, leveraging the complementary strengths of both representations. This fusion enables more robust handling of occluded regions by combining the global context from the SDF with the precise local geometry provided by the point cloud. Extensive experiments on challenging DexYCB and HO3Dv2 benchmarks demonstrate that HOMAE achieves state-of-the-art performance in hand-object pose estimation. We will release our code and model.

cs.CV

CAdam: Confidence-Based Optimization for Online Learning

Modern recommendation systems frequently employ online learning to dynamically update their models with freshly collected data. The most commonly used optimizer for updating neural networks in these contexts is the Adam optimizer, which integrates momentum ($m_t$) and adaptive learning rate ($v_t$). However, the volatile nature of online learning data, characterized by its frequent distribution shifts and presence of noise, poses significant challenges to Adam's standard optimization process: (1) Adam may use outdated momentum and the average of squared gradients, resulting in slower adaptation to distribution changes, and (2) Adam's performance is adversely affected by data noise. To mitigate these issues, we introduce CAdam, a confidence-based optimization strategy that assesses the consistency between the momentum and the gradient for each parameter dimension before deciding on updates. If momentum and gradient are in sync, CAdam proceeds with parameter updates according to Adam's original formulation; if not, it temporarily withholds updates and monitors potential shifts in data distribution in subsequent iterations. This method allows CAdam to distinguish between the true distributional shifts and mere noise, and to adapt more quickly to new data distributions. In various settings with distribution shift or noise, our experiments demonstrate that CAdam surpasses other well-known optimizers, including the original Adam. Furthermore, in large-scale A/B testing within a live recommendation system, CAdam significantly enhances model performance compared to Adam, leading to substantial increases in the system's gross merchandise volume (GMV).

cs.LG

Numerical characterization of the hard Lefschetz classes of dimension two, II: supercritical collections of free divisor classes

For $(n-2)$ free divisor classes on a smooth projective variety of dimension $n$, the product of these free divisor classes induces a Lefschetz type operator acting on the Néron-Severi space or the cohomology group of $(1,1)$ classes. We give a characterization of this kernel space, when the collection of these free divisor classes is supercritical. This resolves Shenfeld-van Handel's open problem in this setting. As consequences, we provide an algebro-geometric proof of the characterization of the extremals of the Alexandrov-Fenchel inequality for a supercritical collection of rational convex polytopes; we also give a characterization of the extremals of the Khovanskii-Teissier inequality given by the intersection numbers of two arbitrary free divisor classes.

math.AG

Channel Estimation for Pinching-Antenna Systems (PASS)

Pinching Antennas (PAs) represent a revolutionary flexible antenna technology that leverages dielectric waveguides and electromagnetic coupling to mitigate large-scale path loss. This letter is the first to explore channel estimation for Pinching-Antenna SyStems (PASS), addressing their uniquely ill-conditioned and underdetermined channel characteristics. In particular, two efficient deep learning-based channel estimators are proposed. 1) PAMoE: This estimator incorporates dynamic padding, feature embedding, fusion, and mixture of experts (MoE) modules, which effectively leverage the positional information of PAs and exploit expert diversity. 2) PAformer: This Transformer-style estimator employs the self-attention mechanism to predict channel coefficients in a per-antenna manner, which offers more flexibility to adaptively deal with dynamic numbers of PAs in practical deployment. Numerical results demonstrate that 1) the proposed deep learning-based channel estimators outperform conventional methods and exhibit excellent zero-shot learning capabilities, and 2) PAMoE delivers higher channel estimation accuracy via MoE specialization, while PAformer natively handles an arbitrary number of PAs, trading self-attention complexity for superior scalability.

cs.IT