Searcharxiv⌕ Search

arXiv subjects

Chao-Kai Wen

Publications and source records attributed to Chao-Kai Wen.

At least 19 recordsLinked to original sources

Rethinking Radiomap Blind Prediction with Limited Environment and Configuration Representations

Radiomap blind prediction infers radiomaps from observable representations of the propagation environment and base station (BS) configuration without field measurements. These representations are inherently incomplete and cannot uniquely determine the target radiomap. Under squared loss, we identify the conditional-mean radiomap as the population-optimal deterministic target and decompose domain risk into target-approximation error and irreducible uncertainty. The train-test risk gap motivates propagation priors as cross-domain guidance, although their partial or simplified forms may bias the attainable predictor. We therefore propose RadioDecomp, which treats a prior-guided predictor as a correctable base and uses deterministic residual refinement to learn its remaining predictable discrepancy. We instantiate RadioDecomp as RadioLSR (LoS-Shadow-Residual). Experiments under cross-configuration and cross-environment settings show that RadioLSR is especially effective for cross-configuration generalization and provides overall gains over a controlled monolithic counterpart under cross-environment generalization.

eess.SP↗

A Graph Foundation Model for Large-Scale MIMO Detection

Large-scale multiple-input multiple-output (MIMO) detection is fundamental to modern wireless networks but constrained by performance-complexity trade-offs. Existing detectors, whether classical or learning-based, often fall short in either scalability or generalizability across heterogeneous scenarios. To overcome these limitations, we introduce a wireless-native graph foundation model (GFM) tailored for large-scale MIMO detection. The proposed GFM employs a physics-informed hybrid architecture, integrating the local correlation extraction of message passing neural networks with the global attention of graph Transformers, encoding the physical interference patterns from the expectation propagation algorithm. Via extensive pre-training, this synergy enables the learning of a general-purpose detection mapping scalable across antenna dimensions and channel conditions. For rapid downstream deployment, parameter-efficient fine-tuning is leveraged to adapt the GFM to specific non-ideal system regimes with minimal overhead. To enhance inference efficiency, a mixture-of-experts mechanism is embedded at downstream deployment to dynamically activate only the necessary sub-modules. Evaluations show that the proposed GFM consistently outperforms classical detectors and advanced data-driven baselines in accuracy, configuration generality, and cross-scenario transferability across various challenging zero-shot and few-shot conditions.

cs.IT↗

Agentic UE-CoMIMO for 6G Terminals: From Virtual Antenna Augmentation to AI-Native Virtualization

End-user-centric collaborative MIMO (UE-CoMIMO) lets nearby devices form a virtual multi-antenna terminal to overcome the antenna limitations of individual user equipment. Extending such cooperation to communication, sensing, computing, and task-relevant information exchange requires a control layer that can interpret user intent, select cooperation mechanisms, and replan as conditions change. This article introduces Agentic UE-CoMIMO, in which device micro-agents, a smartphone or CPE hub agent, and edge/network agents coordinate device participation, relay modes, traffic splitting and duplication, compute placement, semantic-token exchange, and topology reconfiguration. Two system-level scenario studies on creator-centric live streaming and wearable-collaborative blind-spot sensing compare the proposed controller with capability-matched adaptive baselines. The results show that, by anticipating changes and preparing cooperation and fallback actions in advance, agentic control sustains high-quality streaming for longer and maintains blind-spot warnings through device outages. We also discuss the associated standardization, interoperability, trust, and validation challenges.

cs.IT↗

Foundation Models for Wireless Localization: Pretraining, Adaptation, and Utilization

Accurate wireless localization is a key enabler for 6G networks, yet remains challenging under diverse and rapidly changing propagation conditions. Model-based methods degrade when multipath channels are non-resolvable and model mismatches occur, while supervised deep learning demands large labeled datasets and generalizes poorly to new deployments. Inspired by foundation models (FMs) in language and vision, this article presents a unified framework for FM-based wireless localization that learns transferable channel representations from large-scale unlabeled channel state information and adapts to new environments with minimal or even no supervision. We review the fundamentals of FMs, compare the FM paradigm with existing localization approaches, and introduce a three-stage framework spanning large-scale pretraining, localization-oriented fine-tuning, and context-augmented inference, together with the location-aware applications it enables. Ray-tracing-based case studies show improved positioning accuracy and cross-environment generalization. Finally, we present an outlook on key research directions toward AI-native networks for wireless localization.

eess.SP↗

AI/ML Life Cycle Management for Interoperable AI Native RAN

Artificial intelligence (AI) and machine learning (ML) are rapidly becoming integral to the 5G Radio Access Network (RAN), enabling beam management, channel state information (CSI) feedback, positioning, and mobility prediction. However, without a standardized life-cycle management (LCM) framework, challenges such as model drift, vendor lock-in, and limited transparency hinder large-scale deployment. 3GPP Releases 17--20 have progressively introduced AI/ML management and air-interface support, covering model training, validation, deployment, inference, data collection, performance monitoring, applicability assessment, and feature-specific control. Release 20 further extends these capabilities to two-sided CSI compression and inter-vendor model operation. This article reviews the resulting five-block LCM architecture, KPI-driven monitoring mechanisms, and inter-vendor collaboration schemes. We further propose an enhanced LCM framework with detailed interactions across functional blocks and an integrated procedure for reference-model and vendor-model development in two-sided operation, and identify open challenges in resource-efficient monitoring, environment drift detection, intelligent decision-making, and flexible model training. These developments provide a foundation for AI-native transceivers in 6G.

cs.IT↗

Learnware for CSI Feedback: Scene-specific Small Models Can Do Big

Intelligent channel state information (CSI) feedback is essential for realizing the high capacity and spectral efficiency goals of future 6G systems, yet existing deep learning solutions face a trade-off between model generalization and scenario-specific performance. Large neural networks generalize well but incur high computational and tuning costs, while small models excel in particular environments but require repetitive costly end-to-end training for each base station (BS). To address these challenges, we introduce a model repository-based deployment framework in which a centralized AI data center maintains a catalog of scene-specific CSI models. The repository is enhanced with a Learnware-based framework, where each model is associated with a specification including semantic part (network architecture parameters) and statistical part (codeboo-fingerprint embeddings of training-data distributions). A BS submits only its local statistical specifications to retrieve the most relevant pre-trained model, enhancing data privacy by avoiding raw CSI transmission and drastically reducing retrieval latency and communication overhead. We further develop a data-driven search strategy that matches codebook fingerprints to model performance, achieving over 90% selection accuracy. In simulations, our scheme yields 18.8% and 57.7% performance improvements over the General Model in LOS and NLOS scenarios, respectively while reducing local fine-tuning by up to 1000 samples and 100 epochs. This Learnware-based approach minimizes redundant training, maximizes model reuse, and supports rapid,privacy-enhancing deployment of CSI feedback models.

cs.IT↗

CORF-GS: Real-Time Wireless Radiance Field Reconstruction via Coupled Optical-RF Gaussian Splatting

Recent advances in 3D Gaussian Splatting (3DGS)-based wireless radiance field (WRF) reconstruction provide an efficient solution for wireless channel modeling. However, existing WRF reconstruction methods rely on pre-collected observations and offline optimization, and thus struggle to provide real-time channel knowledge. To bridge this gap, we propose CORF-GS, a real-time WRF reconstruction framework that processes sequential optical and radio frequency (RF) keyframes. Specifically, CORF-GS constructs a unified Gaussian representation for optical and RF with shared geometry and modality-specific appearance, allowing high-resolution optical images to provide structural priors for WRF reconstruction. When a new keyframe arrives, CORF-GS first employs optical-guided Gaussian sampling to densify the WRF in under-represented regions. Since light and radio waves may respond differently to the same object surfaces due to wavelength mismatch, relying solely on optical guidance may neglect RF-informative areas. Therefore, CORF-GS performs coupled optical-RF optimization to jointly refine the shared Gaussians. Compared with the existing two-stage training pipelines, this prevents WRF from passively adapting to a frozen optical geometry and encourages the shared Gaussians to adapt to both optical structures and RF power distributions. Simulations show that CORF-GS achieves state-of-the-art RF spectrum synthesis quality and reduces the reconstruction time by $6.4\times$ compared with existing WRF methods.

eess.SP↗

Radar-Aided Near-Field Beam Prediction via Beam Map Learning for XL-MIMO V2I Communications

Near-field beam training in extremely large-scale multiple-input multiple-output (XL-MIMO) vehicle-to-infrastructure (V2I) systems incurs high overhead due to large range-angle codebooks and rapid channel variation. This paper proposes a passive radar-aided framework for near-field beam prediction based on radar-to-beam map learning. By exploiting the spatial correlation between radar observations and communication signals, the proposed method maps radar Bartlett spectra to communication beam maps using a lightweight encoder-decoder convolutional neural network. Gaussian soft supervision is further introduced to preserve beam-space continuity. Simulations on a synchronized Sionna ray tracing radar-communication dataset show that the proposed method consistently improves Top-k accuracy, distance-based accuracy, beam loss, and spectral efficiency.

eess.SP↗

Wireless Imaging for Low-Altitude Surveillance: A New Paradigm for ISAC Networks

The rapid growth of the low-altitude economy calls for integrated sensing and communication (ISAC) networks capable of robust flight monitoring. This article advocates wireless imaging as a unifying sensing paradigm that enables ISAC networks to function as comprehensive low-altitude guardians. We present a hierarchical imaging framework that enhances sensing capability from wide-area snapshot imaging to dynamic trajectory-aware imaging and target-centric fine-grained characterization. At the wide-area level, low-altitude sensing is reformulated as a spatial imaging problem, where distributed base stations and communication users collaboratively construct a holistic view of the aerial space. Building on this foundation, multi-frame dynamic imaging exploits temporal correlations to support robust trajectory tracking and prediction under mobility induced challenges such as occlusions. For security-critical scenarios, the framework further enables fine-grained imaging of flying and hovering uncrewed aerial vehicles, providing detailed characterization beyond conventional point-target abstractions. Additionally, we propose a novel evaluation metric named imaging coverage to examine the sensing fidelity of the proposed framework. Illustrative case studies demonstrate the potential of imaging in ISAC networks to support wide-area monitoring, motion-aware tracking, and fine-grained target analysis.

eess.SP↗

Physics-Informed Path-Parametric Learning for Efficient and Lightweight CSI Feedback

Channel State Information (CSI) feedback is vital for high spectral efficiency in wireless systems, yet high-dimensional CSI introduce significant feedback overhead. Recent deep learning (DL) approaches alleviate this issue by treating CSI as a visual image, but such "black-box" designs often lack interpretability, producing CSI that is not consistent with multipath propagation principles. To address these limitations, this paper proposes HS-PINNnet, a Hierarchical Sensing mechanism assisted Physics-Informed Neural Network for CSI Feedback. Unlike vision-inspired methods, HS-PINNnet integrates a multipath channel model into the network, reformulating high-dimensional CSI reconstruction as low-dimensional multipath parameter estimation (e.g., amplitude, angle). HS-PINNnet features a hierarchical sensing encoder to produce a compact multipath representation, and a heterogeneous decoder for parameter-specific CSI reconstruction, with dedicated branches to estimate different parameters. Moreover, a PCD module adaptively estimates the number of dominant paths in each CSI sample to enhance generalization across diverse environments. A subchannel-wise shared encoding and parallel decoding strategy is further designed to decompose high-dimensional CSI processing into low-dimensional subchannel tasks, reducing training difficulty and improving scalability of HS-PINNnet for future extremely large-scale multiple-input multiple-output (XL-MIMO) systems. Simulation results show that HS-PINNnet outperforms the state-of-the-art under different configurations, achieving a 92.8% reduction in FLOPs and exhibiting two orders of magnitude lower FPGA simulation latency.

eess.SP↗

Digital Twin-Based Channel Generation Toolchain and Foundation Model for Low-Altitude XL-MIMO

The rapid development of the low-altitude economy (LAE) has created growing demand for reliable aerial communication systems. Extremely large-scale multiple-input multiple-output (XL-MIMO) is a promising enabler for such systems due to its high spatial resolution and robust connectivity. However, three-dimensional (3D) mobility together with near-field propagation makes it difficult to obtain dedicated high-fidelity wireless datasets, hindering systematic algorithm development and evaluation. To address this issue, we develop LAETwin-XL, a digital twin (DT)-based toolchain and dataset for XL-MIMO research in LAE scenarios. Built on the Sionna ray-tracing (RT) module, the proposed toolchain simulates near-field and far-field channels with diverse wireless labels for practical environments. Building on this dataset, we further develop a conditional denoising diffusion implicit model (CDDIM)-based generative foundation model that is pretrained to learn transferable XL-MIMO channel representations from incomplete channel observations. Unlike conventional task-specific or foundation models that rely on relatively complete channel inputs, the proposed model can generatively infer informative channel representations from partially observed channels. Experimental results demonstrate that the proposed framework achieves effective zero-shot channel extrapolation performance. Furthermore, using lightweight task heads and limited training data, it enables parameter-efficient transfer to various downstream tasks (e.g., channel estimation, classification, and localization), delivering high accuracy and robustness even under sparse antenna observations. The codes and dataset are available at https://github.com/Lmyxxn/LAETwin-XL.

eess.SP↗

XL-ChannelDiff: An Efficient Diffusion-Based Multi-Domain Near-Field Channel Extrapolation Framework for XL-MIMO Systems

Accurate channel state information (CSI) acquisition is essential for unleashing the performance gains of extremely large-scale multiple-input multiple-output (XL-MIMO) systems. However, in near-field regions, CSI acquisition is much more challenging than in the far field due to the high-dimensional channel representation and spherical wavefront propagation. To address this, in this paper, we propose an efficient multi-domain near-field channel extrapolation framework for XL-MIMO systems. Leveraging the conditional denoising diffusion implicit model (CDDIM), our approach enables accurate channel extrapolation across the antenna, frequency, and spatial domains. Specifically, we design a physics-aware CDDIM backbone that incorporates position-embedded patch tokenization and a mask-guided multi-head attention mechanism, enabling the model to exploit position-dependent channel correlations induced by near-field spherical-wave propagation. To ensure high-fidelity extrapolation, we incorporate a Wasserstein GAN (WGAN) discriminator that provides adversarial supervision to the CDDIM during both the training and reverse sampling phases. Additionally, a RePaint-style refinement scheme is introduced to optimize the sampling trajectory, further boosting extrapolation accuracy. Extensive experiments demonstrate the superiority of the proposed framework, achieving superior extrapolation accuracy and robust generalization across diverse domains, varied configurations, and severe masking conditions.

eess.SP↗

Vision-Based Efficient Joint Trajectory and Channel Tracking in Near-Field XL-MIMO Systems

Accurate joint tracking of mobile users, surrounding scatterers, and dynamic channels is a critical task for sixth-generation (6G) wireless systems, essential for both ensuring high-quality communications and empowering advanced selsing applications such as autonomous driving and immersive extended reality. While extremely large-scale multiple-input multiple-output (XL-MIMO) inherently offers strong support for this task through its high spatial resolution and spectral efficiency, its massive scale of antenna arrays, coupled with near-field propagation characteristics, makes joint trajectory and channel tracking time-consuming and hardware-intensive. To address these challenges, we rethink the problem from a vision-based signal perspective. Specifically, we design a subarray-based partially connected hybrid beamforming (PC-HBF) architecture with a tailored time-multiplexed (TM) mechanism. This effectively compensates for the aperture loss caused by limited radio frequency (RF) chains, generating high-fidelity Cartesian-domain signal images that inherently capture near-field spatial features. Based on this visual representation, we propose an improved CenterNet to perform accurate one-shot path localization, circumventing the path-iterative search required by conventional compressed-sensing-based methods. Building upon this to further improve the accuracy and exploit temporal correlation, a local small-scale orthogonal matching pursuit (OMP) refiner and a lightweight cascaded OMP tracker are developed. Finally, a Hungarian-based trajectory association module is incorporated to maintain track continuity and provide trajectory-level information for environment monitoring. Simulation results show that the proposed framework consistently outperforms representative baselines in position and channel tracking accuracy, especially under low-SNR and limited-hardware conditions.

eess.SP↗

Low-Overhead Receiver Design for Data-Dependent Superimposed Training via Deep Learning

Superimposed pilot (SIP) transmission improves spectral efficiency by eliminating the dedicated pilot overhead required in orthogonal pilot (OP)-based schemes. However, SIP suffers from severe pilot-data coupling, which leads to a critical performance-complexity bottleneck at the receiver. To address this issue, this paper proposes a low-overhead transmission framework that revitalizes data-dependent superimposed training (DDST) with enhanced interference mitigation strategies. First, for quasi-static block-fading channels, an enhanced DDST receiver is developed to achieve non-iterative pilot-data decoupling by exploiting data-dependent algebraic structures. Second, to overcome the sensitivity of conventional DDST to channel variations and symbol misidentification in fast time-varying environments, a mix transmission scheme is developed. By strategically applying DDST to a subset of resource elements, the proposed scheme combines the interference-free transmission property of OP with the zero-pilot-overhead advantage of SIP, thereby improving demapping reliability and interference suppression. Furthermore, under the proposed mix scheme, a Vision Transformer-based neural receiver is designed to capture the orthogonal structure between pilots and perturbation-bearing data, as well as the underlying channel correlations, thereby relaxing the stringent quasi-static assumption required for interference disentanglement. Simulation results demonstrate that the proposed framework achieves significant performance gains in the low-to-medium SNR regime under time-varying channels while providing superior computational efficiency compared with state-of-the-art SIP receivers.

cs.IT↗

Multimodal-NF: A Wireless Dataset for Near-Field Low-Altitude Sensing and Communications

Environment-aware 6G wireless networks demand the deep integration of multimodal and wireless data. However, most existing datasets are confined to 2D terrestrial far-field scenarios, lacking the 3D spatial context and near-field characteristics crucial for low-altitude extremely large-scale multiple-input multiple-output (XL-MIMO) systems. To bridge this gap, this letter introduces Multimodal-NF, a large-scale dataset and specialized generation framework. Operating in the upper midband, it synchronizes high-fidelity near-field channel state information (CSI) and precise wireless labels (e.g., Top-5 beam indices, LoS/NLoS) with comprehensive sensory modalities (RGB images, LiDAR point clouds, and GPS). Crucially, these multimodal priors provide spatial semantics that help reduce the near-field search space and thereby lower the overhead of wireless sensing and communication tasks. Finally, we validate the dataset through representative case studies, demonstrating its utility and effectiveness. The open-source generator and dataset are available at https://lmyxxn.github.io/6GXLMIMODatasets/.

eess.SP↗

Propagation-Consistent Wireless Environment Digital Twin Construction Under Sparse Measurements

Digital twins (DTs) are promising for wireless deployment, optimization, and data generation, but building a propagation-faithful twin from sparse real measurements remains difficult. This paper proposes a wireless environment digital twin (WEDT) construction paradigm that evolves a reconstructed geometric DT into a propagation-consistent wireless environment representation through calibration of a scene-level electromagnetic (EM) property field. Instead of directly fitting link-specific channel responses, the proposed paradigm first constructs a geometry-prior Bayesian channel map (BCM) to convert sparse position-labeled channel state information (CSI) into dense probabilistic supervision with uncertainty estimates. It then embeds the learnable EM property field into differentiable ray tracing (RT) based channel computation, thereby enabling calibration through an explicit propagation chain. Experiments in both public and real-world scenes show that WEDT achieves accurate channel prediction, generalizes to unseen transceiver topologies, and remains effective across different sampling conditions. WEDT also offers utility for material-related environment sensing, more reliable physical-layer planning, and higher-quality synthetic data generation for wireless AI. These results demonstrate the value of the proposed paradigm for propagation-consistent WEDT construction and related wireless applications.

eess.SP↗

AI-Empowered Low-Altitude Economy: Cooperative Sensing With Fixed Wireless Access

The rapid growth of the low-altitude economy has intensified safety concerns arising from unauthorized unmanned aerial vehicles (UAVs), positioning UAV supervision as a key use case in 3GPP. To precisely sense such UAVs with wide coverage and low cost, we leverage fixed wireless access (FWA) customer premises equipment (CPEs), static, densely deployed devices that serve as wireless cameras for the radio environment. We develop an artificial intelligence-empowered two-stage cooperative sensing pipeline that exploits uplink channel state information (CSI) from multiple base station-CPE pairs for UAV detection and localization. In cooperative detection, lightweight CSI features are first individually extracted by neural network, and then adaptively integrated through an attention-based scheme to declare UAV presence. The learned attention scores effectively identify the critical pairs during detection, while facilitating UAV-affected pair selection for subsequent localization. For cooperative localization, neural network initially generates individual estimates and extract CSI features from selected pairs. These estimates, together with features and pair indexes, are fused using a Transformer to produce a precise cooperative estimate. Simulations show that cooperative schemes significantly reduce the missed detection probability to 0.63% and realize a 95%-confidence positioning error of 6.50 m, satisfying 3GPP requirements and showing the potential of FWA-assisted cooperative sensing. Dataset and codes are available on GitHub.

eess.SP↗

Intention-Aware Semantic Agent Communications for AI Glasses

Smart glasses are emerging as a promising interface between humans and artificial intelligence (AI) agents, enabling first-person perception, contextual awareness, and real-time assistance. However, continuous offloading of visual data from wearable devices to cloud-based vision-language models (VLMs) is fundamentally constrained by limited wireless bandwidth and energy resources. This paper proposes an intention-aware semantic agent communication framework for AI glasses, where data transmission is guided by user intention rather than raw pixel fidelity. In the proposed architecture, AI glasses act as an edge semantic agent while a server-side VLM executes high-level cognition and reasoning. The user intention can be inferred by the server-side VLM through the current transmitted content and the historical prompts. Driven by specific user intentions, the glasses adaptively preserve textual content, document layout, or object semantics before transmission. We evaluate three representative scenarios with different lightweight preprocessing tools on the AI glasses. Simulation results demonstrate that intention-aware preprocessing significantly achieves more than 50% bandwidth reduction depending on the current task while maintaining task performance. Moreover, semantic transmission exhibits graceful degradation under low SNRs. The findings demonstrate that aligning communication resources with user intention is essential for robust and efficient wearable AI agent systems.

eess.SP↗