SearcharxivSearch

arXiv subjects

Shijian Gao

Publications and source records attributed to Shijian Gao.

At least 19 recordsLinked to original sources

TeRFS: Temporal-Evolving Radio Field Synthesis

While radio-frequency (RF) field synthesis is fundamental to wireless networking, current approaches remain constrained by static assumptions, leaving them unable to track the rapid multipath reorganization of dynamic scenes. Modeling these transitions requires addressing two coupled challenges: explicit temporal representation and the capture of discrete path lifecycles. To bridge this gap, Temporal-Evolving Radio Field Synthesis (TeRFS) is introduced. TeRFS utilizes an anisotropic spherical Gaussian (ASG) directional basis to represent sparse, sharp angular structures, bound to analytical temporal envelopes that regulate path lifecycles. This formulation induces a mathematical birth-and-death mechanism, enabling individual multipath trajectories to emerge and vanish with temporal precision, a capability beyond the reach of standard smooth interpolation. Evaluations demonstrate that TeRFS outperforms state-of-the-art (SOTA) baselines, achieving an 11.5% reduction in mean squared error (MSE) alongside a 6.9 times training speedup. Even in environments characterized by extreme structural mutation, TeRFS maintains robust tracking of dynamic reorganizations, limiting median absolute error to 1.52 dB and establishing its utility for high-mobility wireless applications. The dataset and code is available at https://github.com/zmydsg/TeRFS.

eess.SP

WiFo-MiSAC: A Wireless Foundation Model for Multimodal Sensing and Communication Integration via Synesthesia of Machines (SoM)

Current learning-based wireless methods struggle with generalization due to the fragmented processing of communication and sensing data. WiFo-MiSAC addresses this as a task-agnostic foundation model that tokenizes heterogeneous signals into a unified space for self-supervised pre-training. A shared-specific disentangled mixture-of-experts (SS-DMoE) architecture is employed to decouple modality-shared and modality-specific representations, facilitating interaction without cross-modal interference. By combining masked reconstruction with contrastive alignment, the model achieves state-of-the-art performance across downstream tasks, including beam prediction and channel estimation. Experimental results demonstrate robust few-shot adaptation and seamless integration of new modalities, positioning WiFo-MiSAC as a scalable backbone for future integrated sensing and communication systems.

eess.SP

WiFo-INR: A Wireless Foundation Model Based on Implicit Neural Representations

Wireless foundation models are emerging as a promising paradigm for AI-native physical-layer design. However, existing methods typically model channel state information (CSI) as image-like discrete tensors with generic token decoders that may struggle to capture complex high-frequency variations efficiently and often produce high-dimensional, size-dependent representations. In this paper, we propose WiFo-INR, an implicit neural representation (INR)-based wireless foundation model that represents CSI as a coordinate-conditioned neural function. A Transformer encoder maps partial or coarse CSI to fixed-dimensional modulation tokens that adapt a SIREN-based decoder, and a compression autoencoder enables quantized CSI feedback. It adopts a two-stage self-supervised pretraining scheme, where mixed masking and denoising improve channel reconstruction and compression-enhanced pretraining enables accurate CSI feedback at low compression ratios. Extensive experiments demonstrate that WiFo-INR learns efficient, compact, and CSI-size-independent implicit wireless representations. Compared with existing foundation models, WiFo-INR improves channel reconstruction and CSI feedback performance while substantially reducing inference latency. It also transfers efficiently to diverse wireless tasks with minimal fine-tuning overhead and achieves zero-shot generalization to unseen CSI sizes.

eess.SP

Stay or Switch: Online Conformal Bayesian Optimization Guided Fluid Antenna Configuration

Fluid antenna systems (FAS) introduce additional spatial degrees of freedom to enable integrated sensing and communication (ISAC) in air-ground networks. However, conventional studies often overlook or simplify the physical overheads and switching costs of FAS. In practice, port switching incurs non-negligible time, during which communication and sensing may continue but with potentially degraded slot-level performance. This leads to two key challenges: (1) the characterization of a slot-level, cost-aware ISAC metric is difficult, and (2) the large port space and accompanying abrupt environmental variations demand more reliable online decision-making. To address these challenges, a cost-aware multi-objective FAS switching problem is formulated, jointly considering slot-level ISAC performance and switching energy. The online conformal Bayesian optimization (OCBO) algorithm is then proposed to learn the unknown gray-box ISAC objectives and calibrate surrogate uncertainty for robust stay-or-switch decisions. Simulation results demonstrate that the proposed cost-aware optimization framework achieves substantially improved long-term ISAC performance compared to existing baselines.

eess.SP

UAV Swarming for Air-Ground ISAC via Cross-Region Cooperation

To serve the volumetric air-ground space, uncrewed aerial vehicles (UAVs) are urgently needed. Yet, relying on them for integrated sensing and communication (ISAC) introduces two key challenges: 1) dynamic and imbalanced ground communication demand, and 2) limited observation diversity for sensing. To address these issues, a cross-region cooperative framework is designed to coordinate UAV swarms. Specifically, a service-driven regional partitioning scheme is proposed to support traffic-aware UAV communication, and an adaptive handshaking mechanism is introduced to improve cooperative sensing accuracy by mitigating residual inter-region phase errors with controlled synchronization overhead. Based on these designs, a region-level multi-agent proximal policy optimization (MAPPO) framework with centralized training and decentralized execution (CTDE) is developed for cross-region cooperative decision-making. Simulation results demonstrate that the proposed method achieves a communication quality-of-service (QoS) of approximately 90% and reduces the Cramér-Rao bound (CRB) by about 45% compared to conventional baselines.

eess.SY

WiFo-M$^2$: Empower Wireless Communications With Plug-and-Play Environment Sensing via Foundation Model

The emerging convergence of next-generation wireless networks and agentic artificial intelligence (AI) is inspiring a new vision: embodied intelligent network entities utilize environmental sensing to refine their physical-layer (PHY) actions. Despite a growing body of preliminary work, prevailing small and task-specific AI models require extensive manual design of data pre-processing, network architecture, and fine-tuning, leaving them tightly coupled to particular PHY actions, system configurations, and deployment scenarios. To address this, we propose a paradigm shift with WiFo-M$^2$, a foundation model that enables environment sensing to be easily integrated into PHY actions, delivering universal performance gains. To extract generalizable out-of-band (OOB) channel-aware features from environment sensing, we introduce ContraSoM, a contrastive pre-training strategy. Once pre-trained, WiFo-M$^2$ infers future OOB channel-aware features from historical sensory data and strengthens feature robustness via modality-specific data augmentation. Experiments show that WiFo-M$^2$ improves the performance of a comprehensive suite of fundamental PHY actions, demonstrating strong generalization to unseen scenarios.

eess.SP

WiFo-2: a generalist foundation model unifies heterogeneous wireless system design

Emerging sixth-generation wireless systems are increasingly heterogeneous, with compatibility across diverse configurations, ubiquitous coverage, and expanded functionalities. Although deep learning has substantially benefited wireless system design, existing approaches are typically trained for specific system settings and scenarios with limited generalizability. Here we present WiFo-2, a space-time-frequency foundation model for unified wireless communications and sensing system design. Pretrained on a heterogeneous dataset of 11.6 billion channel state information (CSI) points, WiFo-2 learns generalized wireless representations across scenarios, configurations, and tasks, and exhibits scaling-law behavior. WiFo-2 achieves reliable and accurate zero-shot channel reconstruction, outperforming fully supervised task-specific models. With only 1% of the training samples required by supervised AI models, WiFo-2 achieves state-of-the-art performance across 9 distinct wireless tasks. A functional hardware prototype further demonstrates its real-world deployability and superior capability across diverse wireless tasks. This work provides a versatile wireless design framework and advances understanding of wireless channels.

eess.SP

Radio Map Updating from Streaming Spectrum Measurements via Memory-Based Online Gaussian Processes

Radio maps, which estimate spatial radio-frequency characteristics from spectrum measurements, are essential for applications such as spectrum management and network planning. With the continuous arrival of spectrum measurements, conventional batch processing methods for radio map reconstruction become computationally prohibitive, as they require reprocessing all accumulated measurements for each radio map update. To address this, we propose a memory-based online sparse variational Gaussian process (M-OSVGP) method that efficiently updates radio maps from streaming spectrum measurements. Our method employs sparse variational inference and updates the posterior online by minimizing a hybrid objective that integrates newly received measurements and a memory subset of previous ones to mitigate catastrophic forgetting. To further improve posterior approximation as measurements accumulate over spatially diverse regions, we extend M-OSVGP with a grid-assisted online inducing point selection (GOIPS) algorithm. GOIPS dynamically adapts the number and locations of inducing points based on measurement density and spatial correlation, providing a more informative inducing set while maintaining computational efficiency. Extensive simulations demonstrate the effectiveness of our proposed methods in reconstruction accuracy, computational efficiency, and uncertainty quantification, compared to existing batch and online baselines across various scenarios.

eess.SP

FARM: Foundational Aerial Radio Map for Intelligent Low-Altitude Networking

Precise aerial radio environment characterization is vital for low-altitude airspace planning. However, existing datasets and construction methods lack the high-resolution granularity required for complex aerial spaces, particularly failing to capture spatial variations across both horizontal and vertical dimensions. To address these gaps, this paper introduces FARM, a pioneering foundation model for unified aerial radio map (ARM) construction. FARM is supported by our newly curated, high-granularity full-domain ARM dataset, which features multi-band and multi-antenna configurations, effectively filling a critical void in comprehensive low-altitude radio data. Structurally, FARM leverages a masked autoencoder to extract deep latent representations of the aerial radio environment, which subsequently guide a diffusion-based decoder to synthesize high-fidelity signal distributions through only a few iterative refinement steps. Benefiting from this design, the architecture seamlessly accommodates both condition-based and condition-free ARM construction, providing robust support for diverse signal and environmental priors. Extensive experiments demonstrate that FARM significantly outperforms state-of-the-art benchmarks while exhibiting strong cross-scenario generalization. Crucially, we validate the transferability of FARM on a real-world dataset collected from field tests, proving its robust deployment capability. Ultimately, FARM serves as a foundational infrastructure for the low-altitude economy by enabling autonomous aerial logistics and intelligent urban networking.

eess.SP

Sensing-Native Over-the-Air Federated Learning

Over-the-air federated learning (FL) leverages the superposition property of multiple-access channels to enable communication-efficient distributed model training. Existing integrated sensing, communication, and computation (ISCC)-enabled over-the-air FL systems typically require dedicated resources for the sensing module, inevitably compromising FL performance due to resource competition. In this paper, we propose a sensing-native over-the-air FL framework that explores built-in distributed wireless sensing capability with zero overhead per model aggregation. Specifically, the high-dimensional local gradient signals possessing favorable autocorrelation property are concurrently leveraged for target distance estimation, while the gradient statistics already required for over-the-air FL serve as a ready-made gateway to deliver locally-sensed results to the edge server for cooperative localization. To combat inter-device interference, channel fading, and communication noise, we put forth a robust trilateration-based target positioning method building upon an efficient matched-filtering-based distance estimation. Then, by explicitly characterizing the impact of imperfect model aggregation and noisy gradient-statistics transmission on the sensing-native over-the-air FL convergence, we develop a statistics-aware communication-learning co-design approach. We first derive the closed-form optimal power budgets allocated to local gradients and their statistics, based on which an efficient successive convex approximation method is proposed for receiver beamforming optimization. Simulation results show that the proposed framework simultaneously achieves superior learning and sensing performance compared to representative baselines.

eess.SP

Active Perception for Radio Map Reconstruction in Uncharted 3D Air-Ground Environments

Radio maps provide the essential foundation for low altitude networking systems. Unlike terrestrial radio maps that are typically generated via drive test measurements, mapping the air-ground environment requires the deployment of unmanned aerial vehicles (UAVs). This shift introduces two formidable challenges in uncharted 3D scenarios. First, sparse radio measurements and incomplete geometric observations hinder accurate reconstruction. Second, the large 3D action space and strict power constraints from high spectrum scanner energy consumption make informative exploration difficult. To address these issues, this paper proposes 3D uncertainty aware radio active mapping (3D-URAM), a closed loop active perception framework that decouples the mapping process into two offline trained stages. In Stage I, a Bayesian UNet is developed to recover radio maps from sparse measurements and partial geometry while providing calibrated predictive uncertainty. In Stage II, a dynamic probabilistic roadmap and a transformer based waypoint selection policy trained via proximal policy optimization maximize long horizon uncertainty reduction under travel budgets. Experimental results demonstrate that 3D-URAM reduces reconstruction error by over 50% compared to representative baselines. Real-world field tests within a 300mx200mx100m space also validate the potential of active radio map reconstruction.

eess.SP

Learn to Access and Backhaul the Sky: Multi-Scale Radio Map Guided Multi-UAV Cooperation

Driven by the emerging low-altitude economy, uncrewed aerial vehicle (UAV) swarms offer flexible integrated air-ground access and backhaul. However, providing seamless connectivity is difficult due to the interdependent dynamics of user mobility and building blockages in these 3D scenarios. These factors create rapidly shifting bottlenecks in end-to-end paths. Furthermore, the multi-dimensional nature of joint control limits the effectiveness of traditional heuristics. To address these challenges, a \textbf{\underline{M}}ulti-Scale \textbf{\underline{R}}adio \textbf{\underline{M}}ap-\textbf{\underline{G}}uided (MRMG) framework is proposed. The MRMG framework handles heterogeneous dynamics by integrating three distinct levels of radio information: global-level maps provide regional coverage insights, local-level maps capture neighborhood-scale service conditions, and link-level maps characterize high-resolution channel features. This design effectively decouples macro-movement from micro-link adaptation. To yield long-term performance improvements, A multi-agent reinforcement learning (MARL) controller learns cooperative policies for UAV movement, next-hop selection, and transmit-power control. Simulation results show that the MRMG framework not only improves network throughput but also significantly bolsters cell-edge service, nearly doubling the 5th-percentile user rate.

eess.SP

Sequential Task Assignment and Resource Allocation in V2X-Enabled Mobile Edge Computing

Nowadays, the convergence of mobile edge computing (MEC) and vehicular networks has emerged as a vital enabler for the ever-increasing intelligent onboard applications. This paper proposes a multi-tier task offloading mechanism for MEC-enabled vehicular networks leveraging vehicle-to-everything (V2X) communications. The study focuses on applications with sequential subtasks and explores the collaboration of two tiers. In the Vehicle Tier, the requesting vehicle (RV)-service vehicle (SV) matching scheme and the inter-vehicle collaborative computation are studied, with joint optimization of task offloading decision, communication, and computing resource allocation to minimize energy consumption while satisfying delay requirements. In the Roadside Unit (RSU) Tier, collaboration among RSUs is investigated to further address multi-access issues of uplink subchannels and computing resources for serving unmatched RVs. To tackle this intricate problem, a layered optimization framework is first proposed to obtain task offloading decisions and optimal continuous resource allocation, after which a subchannel allocation scheme is designed to recover the discrete solution with low complexity. Extensive experiments are conducted to demonstrate that the proposed method reduces average energy consumption by at least 15% compared with recent utility maximization and energy cost minimization benchmarks under varying task delay requirements and vehicle scales.

cs.NI

Dynamic Task and Resource Scheduling Towards Green Space-Air-Ground-Sea Integrated Network

In the context of 6G ubiquitous connectivity, the space-air-ground-sea integrated network (SAGSIN) emerges as a new paradigm to provide critical services for resource-limited ocean environments. To realize this paradigm efficiently, we propose an innovative dynamic task and resource scheduling approach for green SAGSIN that delivers computing support for vessels while minimizing overall task execution delay. To address the challenge of multi-layer task scheduling, a layer-wise task offloading algorithm is developed specifically for SAGSIN. It adapts to real-time, multi-dimensional system dynamics and integrates an anticipatory handover strategy that adaptively controls the amount of data offloaded to the satellite, thereby preventing post-handover congestion while improving satellite resource utilization. Furthermore, the bandwidth allocation of uncrewed aerial vehicles and base station, UAV trajectories, and computing resource allocation are jointly optimized to enhance connectivity among low-altitude devices and facilitate demand-driven resource allocation for green network development. Simulation results verify that the proposed method better adapts to dynamic system resources and achieves at least a 23% reduction in average task delay compared with benchmarks.

cs.NI

Scalable Multimodal Beam Alignment in V2X: An Anti-Imbalance Graph Learning Approach

Efficient beam alignment is fundamental to high-throughput and reliable connectivity in Vehicle-to-Everything (V2X) systems. However, conventional beam management in dynamic vehicular topologies incurs prohibitive alignment overhead and struggles to maintain robust links under rapid mobility. To overcome these challenges, this paper proposes a distributed multimodal graph beam alignment (GBA) framework. The core innovation lies in leveraging onboard multimodal sensing data to predict implicit feedback while employing graph neural networks to coordinate multi-user alignment, thereby jointly enhancing scalability and drastically reducing overhead. The architecture adopts a dual-network design with GBA-RSU and GBA-Vehicle units, optimized through a hybrid strategy of centralized learning and federated learning (FL) to balance global performance with local privacy. Furthermore, a dedicated data augmentation (DA) scheme is introduced to address multimodal data imbalance issues in vehicular networks. Negative augmentation applies dominant modality dropout to bolster robustness, while positive augmentation generates underrepresented samples to mitigate label imbalance. Numerical results demonstrate that GBA maintains a competitive sum rate on par with high-resolution codebook-based feedback yet reduces beam alignment overhead by over 90\% and scales efficiently in mobile scenarios. Notably, integrating DA enables GBA to consistently outperform state-of-the-art FL-based alignment benchmarks, with particularly pronounced gains under severe label and modality imbalance, establishing a practical solution for V2X beam management.

eess.SP

Grey-Box Bayesian Optimization for ISAC in Fluid-Antenna Assisted Air-Ground Network

Fluid antenna systems (FAS) provide extra position agile spatial diversity for integrated sensing and communication (ISAC), by jointly optimizing the port selection and precoding. However, this optimization is challenging in air ground networks due to the intricate dual objective Pareto frontier, complex self-interference, and prohibitive channel state information overhead. To overcome these bottlenecks, this work proposes a novel grey box multi objective Bayesian optimization framework to address the joint design of discrete port selection and ISAC precoding. Unlike black box methods, this architecture explicitly leverages known physical system models to learn unknown channel constituents, dramatically reducing sample complexity. To navigate high dimensional combinatorial spaces, an adaptive trust region mechanism powered by expected hypervolume improvement (EHI) acquisition is implemented. Furthermore, the framework incorporates a spatio-temporal tracking strategy to handle the continuous mobility of users and targets, robustly capturing the drifting optimum in time varying environments. Simulations demonstrate that this framework achieves significantly faster convergence and discovers superior Pareto optimal configurations, validating its efficiency for dynamic real time FAS-ISAC deployments.

eess.SP

Score-Based Conditional Flow Models for MIMO Receiver Design with Superimposed Pilots

Accurate channel state information (CSI) is vital for multiple-input multiple-output (MIMO) systems. However, superimposed pilots (SIP), which reduce overhead, introduce severe pilot contamination and data interference, complicating joint channel estimation and data detection. This paper proposes a conditional flow matching receiver (CFM-Rx), an unsupervised generative framework that learns directly from received signals, eliminating the need for labeled data and improving adaptability across diverse system settings. By leveraging flow-based generative modeling, CFM-Rx enables deterministic, low-latency inference and exploits model invertibility to capture the bidirectional nature of signal propagation. This framework unifies flow matching with score-based diffusion modeling via a moment-consistent ordinary differential equation (ODE), replacing stochastic differential equation (SDE) sampling with a deterministic and efficient process. Furthermore, it integrates receiver-side priors to ensure stable, data-consistent inference. Extensive simulation results across various MIMO configurations demonstrate that CFM-Rx consistently outperforms conventional estimators and state-of-the-art data-driven receivers, achieving notable gains in channel estimation accuracy and symbol detection robustness, particularly under severe pilot contamination.

eess.SP

Transfer to Sky: Unveil Low-Altitude Route-Level Radio Maps via Ground Crowdsourced Data

The expansion of the low-altitude economy is contingent on reliable cellular connectivity for unmanned aerial vehicles (UAVs). A key challenge in pre-flight planning is predicting communication link quality along proposed and pre-defined routes, a task hampered by sparse measurements that render existing radio map methods ineffective. This paper introduces a transfer learning framework for high-fidelity route-level radio map prediction. Our key insight is to leverage abundant crowdsourced ground signals as auxiliary supervision. To bridge the significant domain gap between ground and aerial data and address spatial sparsity, our framework learns general propagation priors from simulation, performs adversarial alignment of the feature spaces, and is fine-tuned on limited real UAV measurements. Extensive experiments on a real-world dataset from Meituan show that our method achieves over 50% higher accuracy in predicting Route RSRP compared to state-of-the-art baselines.

eess.SP