Searcharxiv⌕ Search

arXiv subjects

Jianhua Zhang

Publications and source records attributed to Jianhua Zhang.

At least 19 recordsLinked to original sources

A Tutorial on Learning-Based Radio Map Construction: Data, Paradigms, and Physics-Awareness

Radio maps (RMs) provide the digital representation of the wireless propagation environment, mapping complex geographical and topological boundary conditions to critical spatial-spectral metrics that range from received signal strength to full channel state information matrices. The integration of artificial intelligence into next generation wireless networks further necessitates the accurate construction of RMs as a foundational prerequisite for electromagnetic digital twins. This paper presents a comprehensive survey of learning-based RM construction, systematically addressing three intertwined dimensions: data, paradigms, and physics-awareness. From the data perspective, we review physical measurement campaigns, ray tracing simulation engines, and publicly available benchmark datasets, identifying their respective strengths and fundamental limitations. From the paradigm perspective, we establish a core taxonomy that categorizes RM construction into source-aware forward prediction and \added{source-agnostic} inverse reconstruction, and examine five principal neural architecture families spanning convolutional neural networks, vision transformers, graph neural networks, generative adversarial networks, and diffusion models. We further survey optics-inspired methods adapted from neural radiance fields and 3D Gaussian splatting for continuous wireless radiation field modeling. From the physics-awareness perspective, we introduce a three-level integration framework encompassing data-level feature engineering, loss-level partial differential equation regularization, and \added{architecture-level} structural isomorphism. Open challenges including foundation model development, physical hallucination detection, and amortized inference for real-time deployment are discussed to outline future research directions.

eess.SY↗

Measurement-Based FR3 Urban Macrocell Channel Characterization and Coverage Analysis

The 6--18~GHz upper mid-band is a promising spectrum range for future sixth-generation (6G) networks due to its favorable balance between available bandwidth and propagation capability, while its same-site coverage performance remains a key deployment concern. This paper presents a wideband channel measurement campaign at 13 spatially aligned frequency points from 6 to 18~GHz in an urban macrocell (UMa) environment, and investigates the frequency evolution of channel characteristics and same-site coverage capability. The results show that path loss generally increases with frequency, while the higher-frequency channels tend to exhibit smaller root-mean-square (RMS) delay spread (DS) and larger PDP-based $K$-factors. Based on the measured path-loss characteristics, the additional link-budget requirement and service-aware same-site coverage are further evaluated. Under the baseline service configuration, the 80\% coverage distance $d_{80}$ decreases from approximately 129~m at 6~GHz to 58~m at 18~GHz. Further analysis quantifies the impacts of service rate, resource allocation, interference, and link-budget degradation on coverage, providing measurement-based insights for same-site deployment across the 6--18~GHz band.

eess.SP↗

Reassessing 3GPP NR CSI Codebook Structures in Near-Field Channels: Finite-Feedback Multilayer Precoding and Design Insights

3GPP TR 38.901 Rel-19 introduces antenna-element-level spherical-wave modeling, while NR Type-I and enhanced Type-II (eType-II) CSI codebooks continue to use plane-wave DFT beams. Whether this mismatch materially degrades finite-feedback multilayer precoding in standardized multipath channels remains unclear. To isolate its impact, we evaluate both codebooks over strictly paired far-field (FF) and near-field (NF) 3GPP channels that share user locations, multipath parameters, polarization, and link budgets and differ only in their wavefront models. Simulations cover Rank 1-4 transmission in 7-GHz UMi and 24-GHz InH-linear scenarios. We find no systematic FF/NF shift in singular-mode gains or equal-power SVD (SVD-EP) rates. Instead, spherical-wave phases reorder multipath projections onto plane-wave candidates and thereby alter codeword selection. For Rank-4 InH-linear users at 0.1 normalized Rayleigh distance, given FF-selected Type-I and eType-II codewords incur median direct mismatch losses of 2.28% and 5.34% on the NF channel, respectively; codebook reselection identifies better-matched codewords and reduces these losses to 0.770% and 3.28%. Nested candidate-set comparisons further show that relaxing Type-I interlayer constraints improves the SVD-EP-normalized rate by 20.7 percentage points, whereas finite-range sampling adds only 0.579 points. These results support prioritizing multilayer multibeam representation in large-aperture NR CSI codebooks, with range states providing complementary refinement.

eess.SP↗

Exact Degrees of Freedom of Spatially Sparse MIMO Channels Without Prior CSI

We characterize the degree of freedom (DoF) of a point-to-point blockwise memoryless channel without prior channel state information (CSI), with a fixed number $K$ of propagation paths, where the transmitter (Tx) and the receiver (Rx) are equipped with nonuniform linear arrays (NULAs) of $N_t$ and $N_r$ antennas, respectively. The positions of array elements are fixed, known, pairwise distinct, and need not be equally spaced. The uniform linear array (ULA) is a special case. In each block of length $T$, the continuous angles of arrival (AoAs), angles of departure (AoDs), and independent complex Gaussian path gains are redrawn. Both Tx and Rx know the state distributions but are not given the current realizations before transmission. The receiver may estimate the channel from reference signals or decode without explicit channel estimation, with reference symbols counted in $T$ and their energy counted against the power constraint. Under the aforementioned model, we show that the DoF is $1-\frac{1}{T}$ for $K=1$, and $K(1-\frac{3}{2T})$ for $K \geq 2$, when $N_r\ge K+1$, $N_t\ge\max\{K,2\}$, and $T\ge K$. The analytical results are further demonstrated by their applications to the DoF tradeoff analysis in integrated sensing and communication (ISAC). For more general array structures, an achievability result is established, while the converse remains open in general.

cs.IT↗

Computing at Sea: Floating and Offshore Data Centres as a Pathway to Sustainable AI Infrastructure

The rapid expansion of artificial intelligence is transforming data centres into one of the world's fastest-growing sources of electricity demand. As AI systems scale in size and capability, the physical infrastructure supporting computation is approaching critical limits in energy availability, cooling capacity, land use, freshwater consumption, and carbon management. Conventional land-based data centres are increasingly constrained by urban land competition, grid congestion, environmental pressures, and lengthy permitting processes, raising fundamental questions about where future computing infrastructure can sustainably exist. This article examines floating and offshore data centres as an emerging alternative model for digital infrastructure. By relocating computation to marine environments, offshore systems can exploit the ocean's natural cooling capacity, reduce freshwater dependence, and enable direct integration with offshore renewable energy resources such as wind, wave, and tidal power. Early deployments have demonstrated the potential for significantly improved energy efficiency and operational reliability compared with conventional facilities, while also opening new possibilities for distributed and resilient computing architectures. The article explores how offshore computing may reshape the future relationship between electrification, renewable energy, and large-scale AI infrastructure. It analyses the opportunities and trade-offs associated with marine deployment, including environmental impacts, engineering design challenges, economic feasibility, and regulatory governance. Rather than treating offshore data centres as experimental novelties, the article presents them as part of a broader systems-level transition in how society may power, cool, and sustain the next generation of computational growth.

eess.SY↗

PICANet: Physics-Informed Cascaded Asymmetric Network for Infrared Small Target Detection

Infrared small target detection (ISTD) is an important research direction in image processing. However, existing methods are limited by severe background noise propagation and target degradation in high-level semantic features. To address these limitations, this paper proposes a plug-and-play physics-informed cascaded asymmetric network, named PICANet. Specifically, we construct a hierarchical prior decoupling module to explicitly extract low-level and high-level physical information, thereby characterizing target features at different levels rather than relying solely on convolutional extraction. Furthermore, a dual-prior interactive fusion module is developed to dynamically refine target representations while suppressing complex background clutter. Unlike previous work, a multi-level cross-feature attention module with the cascaded asymmetric mechanism is introduced to achieve precise alignment between high-level semantics and low-level spatial details. Extensive experiments demonstrate that the proposed PICANet outperforms state-of-the-art ISTD methods, showing satisfactory detection accuracy even against complex backgrounds. Our code is available at https://github.com/xianchaoxiu/PICANet.

cs.CV↗

Site-specific Channel Modeling Based on Remote-Sensing Maps for 6G Space--Air--Ground Digital Twins

Site-specific channel models are essential for wireless digital twins of 6G space--air--ground communication systems. However, 3D maps are difficult to obtain over wide areas, which limits large-area site-specific channel modeling. To address this issue, this paper proposes a remote-sensing-based augmented ray-tracing channel modeling framework. The framework comprises a deterministic RT branch, a measurement-statistical branch, and an RT augmentation branch. To overcome the difficulty of acquiring large-area 3D maps, the deterministic RT branch reconstructs a 3D RT scene from satellite remote-sensing imagery and calibrates its electromagnetic material parameters using measured path loss. To provide the statistical parameters required for RT augmentation, the measurement-statistical branch establishes the marginal distributions and interparameter dependence models of the channel parameters. Specifically, a wideband UAV channel measurement campaign is conducted at 4.60 GHz, and a proposed multipath estimation method estimates the complex amplitudes, delays, and Doppler shifts of the measured multipath. To bridge the gap between RT predictions and measurements, the RT augmentation branch organizes the RT multipath into LoS, LoS-tail, and NLoS components, generates additional short-delay LoS-tail paths, and reallocates the component and path powers according to the measurement-derived statistics while preserving the total RT received power. The validation results show that the proposed framework reduces the path loss RMSE from 5.45 to 4.35 dB and, relative to calibrated RT, decreases the RMS delay spread and normalized Doppler spread RMSEs by 53.03 and 26.48, respectively. The proposed framework provides a site-specific channel modeling approach for 6G space--air--ground digital-twin studies.

eess.SP↗

Electromagnetic World Model for 6G: A Unified Framework for Joint Environment Reconstruction and Channel Prediction

The integration of sensing, communication, and intelligence is becoming a key enabler for sixth generation (6G) wireless systems, where intelligent terminals are expected to simultaneously support efficient link establishment and reliable environmental sensing. However, existing studies mainly exploit sensing information or communication information to address a single task, such as channel prediction or environment reconstruction. Motivated by the shared dependence of optical and radio-frequency signals on the surrounding environment, we propose the electromagnetic world model (EMWM), the first unified framework for joint environment reconstruction and channel prediction. EMWM learns a common electromagnetic representation with the potential to provide a modeling foundation for 6G tasks. Specifically, partial channel state information (CSI) and multi-view red-green-blue (RGB) images are encoded into CSI and visual tokens and jointly processed by a hierarchical world-model backbone with local and global aggregation. Based on the learned representation, a mixture-of-experts (MoE)-based CSI prediction head reconstructs the complete CSI, while a depth prediction head estimates multi-view depth maps that are further converted into three-dimensional (3D) point clouds. Moreover, a large-scale multi-modal dataset is constructed based on a campus digital twin. Experimental results show that EMWM outperforms conventional neural network and large language model (LLM) baselines in both CSI prediction and environment reconstruction, achieving a squared generalized cosine similarity (SGCS) of 0.9699 for CSI prediction while demonstrating robustness across different signal-to-noise ratio (SNR) conditions and zero-shot generalization at 28 GHz.

eess.SP↗

Revisiting Shannon's Source Coding Theorem with Distributional Uncertainty under the Nonlinear Expectation Theory

In classical information theory, a source is modeled by a single, precisely known probability distribution. However, in the increasingly complex communication networks full of unanticipated, nonstationary, and heterogeneous random events, the assumption of precise and well-defined probability distributions to describe random variables appears somewhat idealized. Therefore, it is important to characterize the uncertainty of distributions of source messages, subject to relaxing the assumption of deterministic probability models for analyzing information sources in information theory. Based on the nonlinear expectation theory, a novel axiomatical system that extends classical probability theory, this paper investigates the information sources whose distributions themselves are uncertain, and refers to them as uncertain-distribution sources. We generalize the fundamental concept information entropy to nonlinear information entropy, which describes the measurement of the amount of information contained in a uncertain-distribution source. By using the strong law of large numbers under sublinear expectation, we establish a nonlinear source coding theorem, which not only shows that the nonlinear information entropy is the upper bound for the infimum of achievable coding rate of uncertain-distribution sources under the maximum error probability criterion, but also determines a cluster point of the coding rate of uncertain-distribution sources under the minimum error probability criterion. Our findings reveal that the introduction of nonlinear expectation theory allows for a more comprehensive understanding of information sources.

cs.IT↗

WiWorld-RealData: A Real-World Multi-Modal Dataset for 6G Wireless World Models

As sixth-generation wireless systems evolve from reliable connectivity toward environment intelligence, wireless world models aim to learn how physical environments and user states affect wireless propagation, requiring real-world data with explicit correspondences between channel responses and environment observations. However, existing channel-environment datasets are predominantly simulation-based or designed for specific communication tasks, limiting their support for general environment-channel relationship learning. To address this gap, we construct WiWorld-RealData, a real-world multi-band channel and multi-modal environment sensing dataset for 6G wireless world model research. It provides synchronized channel impulse responses measured at 3.7 and 6.775 GHz together with multi-view and panoramic images, light detection and ranging point clouds, millimeter-wave radar observations, and global navigation satellite system trajectories. Unified timestamps, sample identifiers, and metadata establish sample-level correspondences across these heterogeneous modalities. The overall measurement campaign produced approximately 10 TB of data, while the current public release provides aligned channel-environment samples from a representative continuous outdoor route. A path-loss prediction case study further validates the dataset using a continuous test route segment, achieving a mean absolute error of 2.02 dB and a root mean square error of 2.69 dB under few-shot adaptation. WiWorld-RealData supports cross-band propagation analysis, environment-aware channel modeling, wireless digital twins, and channel foundation model research. The dataset is available at https://scc.bupt.edu.cn/dataset-manage/datasets/44 and https://doi.org/10.57760/sciencedb.40663.

eess.SP↗

New Mid-Band (FR3, 6-24 GHz) XL-MIMO for 6G: Channel Modeling, Algorithm Evaluation, and Field Trials

The new mid-band (FR3, 6-24 GHz) spectrum is expected to play an important role in future 6G networks by providing a favorable balance among coverage, capacity, and deployment feasibility. Meanwhile, extremely large-scale multiple-input multiple-output (XL-MIMO) has emerged as a key enabling technology to exploit the propagation and spatial multiplexing potential of these frequency bands. Firstly, this paper provides a systematic review of spectrum allocation and standardization activities for new mid-band spectrum, together with the 6G spectrum planning strategies of countries and regions. Secondly, the wideband massive MIMO channel sounder is also introduced, which is specially developed for channel measurements of new mid-band with over a thousand elements. Thirdly, propagation characteristics and channel modeling approaches of four representative XL-MIMO architectures, including co-located, cell-free, and intelligent XL-MIMO, are comprehensively reviewed and analyzed, with particular emphasis on near-field propagation, spatial non-stationarity, and capacity performance. Then, recent advances in channel estimation, beamforming, and artificial-intelligence-assisted signal processing are summarized. In addition, the performance of new mid-band XL-MIMO systems equipped with 1536 and 768 antenna elements is comparatively evaluated. Finally, real communication environment prototype system field trials conducted in the Upper 6 GHz (U6GHz) band are used to investigate practical system performance under realistic deployment conditions. The results indicate that the target signal-to-noise ratio is a critical factor affecting XL-MIMO performance in the U6GHz band.

eess.SP↗

Machines that know they are aging: a framework for hardware-aware autonomous intelligence

Autonomous systems inevitably age, yet their artificial intelligence typically assumes hardware remains in its original condition. Batteries degrade, sensors drift, processors accumulate timing errors, and memory reliability declines, creating a growing mismatch between assumed and actual capability. This can lead to agnostic collapse, where mission failure arises from accumulated hardware degradation rather than a single component fault. We propose Aging-Aware Autonomous Intelligence (AAAI), a framework that integrates hardware health directly into reasoning, planning, and mission execution. AAAI is built on three pillars: hardware self-awareness, which continuously estimates the health of power, sensing, memory, and computation subsystems using physics-of-failure models; self-adaptive reasoning, which adjusts inference complexity, planning horizon, and task priorities according to remaining hardware capability; and survival-centric intelligence, which allocates remaining operational life across mission objectives through performance optimization, resource conservation, and graceful degradation. Rather than introducing new hardware, AAAI unifies prognostics, lifecycle management, and hardware-aware computing into a closed-loop cognitive architecture. We argue that such integration is essential for autonomous systems operating in inaccessible or safety-critical environments, including space missions, marine robotics, and implantable medical devices. By enabling machines to recognize and respond to their own aging, AAAI improves resilience, extends operational lifetime, and supports safer, more graceful mission completion.

cs.RO↗

MVLA-GR: A Phase-Free Multipath-Based Geometry Reconstruction Method via Multi-View Likelihood Accumulation for ISAC

Integrated sensing and communication (ISAC) enables wireless systems to reuse communication signals for environmental sensing, where reconstructing the geometry of surrounding objects is a representative sensing task. However, many conventional methods rely on coherent processing and require accurate phase information, which is often hard to guarantee in practical communication systems, particularly at high carrier frequencies. To address this problem, this paper proposes a Multi-View Likelihood Accumulation Geometry Reconstruction (MVLA-GR) method based on channel impulse response (CIR) measurements, which uses only delay and power observations without requiring phase information. The method extracts dominant multipath components from each observation, and for each candidate spatial location, accumulates components across views whose propagation distances match the location as supporting evidence. A soft distance-matching kernel is introduced to tolerate range estimation errors and viewpoint-dependent scattering migration, and the received power of each component is used as a reliability weight. A joint thresholding strategy combining response magnitude and angular support continuity then converts the continuous support map into a binary geometry estimate. Ray-tracing simulations on canonical and complex targets, as well as real-world vehicle measurements at 36 GHz, demonstrate that MVLA-GR can effectively recover target geometry, providing a low-complexity phase-free solution for ISAC.

eess.SP↗

WEKP-PLP: A Wireless Environment Knowledge Pool-Enhanced Path Loss Prediction Framework for Shore-to-Ship Communication

Accurate shore-to-ship path loss prediction is essential for maritime mobile communication systems, but it remains challenging because dynamic sea surface conditions can alter reflection paths and multipath effects, introducing uncertainty into the observed path loss. This paper proposes a wireless environment knowledge pool-enhanced path loss prediction (WEKP-PLP) method for point prediction and interval characterization of shore-to-ship path loss. The proposed method first constructs a wireless environment knowledge pool (WEKP) from ray tracing (RT) simulations under different wind speeds, temperatures, salinities, frequencies, antenna heights, and propagation distances. The WEKP learns the mapping from environmental and system parameters to path loss residual quantiles and provides a prior that contains both the median prediction and the associated prediction interval. To adapt this prior to real scenarios, a small number of measurement samples are used to learn the residual between the WEKP median prediction and the measured path loss through Gaussian process regression (GPR). The experimental results show that the WEKP-PLP provides accurate point prediction and compact intervals with reliable coverage. The proposed method achieves an MAE of 1.03 dB and an RMSE of 1.54 dB. The ablation results further confirm that both the WEKP prior and the residual correction using a small number of measurement samples are necessary for accurate and reliable shore-to-ship path loss prediction.

eess.SP↗

What the Waveform Knows: Transparent-first Speech and Audio Intelligence with Caption Studio

Caption Studio is a transparency-first speech and audio intelligence platform that transforms spoken audio and video into structured, searchable content through automated transcription, speaker diarization, speech analytics, signal-level audio analysis, and subtitle generation. The system is built on a FastAPI backend with a real-time dashboard and adopts a three-layer architecture comprising (i) a transcription and diarization core based on Whisper-class automatic speech recognition and pyannote speaker diarization, (ii) an audio intelligence layer that extracts acoustic and linguistic features, including waveforms, spectrograms, pitch, speaking rate, silence, filler-word frequency, and sentiment, directly from the audio signal, and (iii) an integration layer that supports data export and downstream workflow integration. A principal contribution of this work is the transparency-first framework, in which every reported metric is explicitly identified as measured, derived, or unavailable, thereby improving the traceability, interpretability, and reliability of speech analytics. The paper presents the system architecture, benchmarking methodology, explainability and uncertainty framework, and key considerations for enterprise-scale deployment.

cs.SD↗

LFM: Leveraging Foundation Models for Source-Free Universal Domain Adaptation

Source-free universal domain adaptation (SF-UniDA) adapts a pre-trained source model to an unlabeled target domain under both covariate and label shifts, without access to source data. However, existing SF-UniDA methods rely on inefficient techniques such as threshold tuning and clustering. Foundation models (FMs), known for their generalization and zero-shot capabilities, remain underexplored in SF-UniDA. In this paper, we propose a framework that leverages foundation models (LFM) for SF-UniDA. We use a vision-language model (VLM) to compute similarities between target samples and text labels, including those for unknown classes generated by prompting a large language model. The label shift type is determined by analyzing the coefficient of variation of a similarity-based sample-level score. Unknown samples are identified using a binary Gaussian mixture model fitted to another similarity-based metric. Under a consensus strategy, the pseudo-labels generated by the VLM are refined by the target model initialized with the pre-trained source model, integrating knowledge from both the source domain and foundation models. Finally, these refined pseudo-labels are used to train the target model. Extensive experiments across all possible label shifts and multiple benchmarks demonstrate the effectiveness and superiority of our proposed LFM framework. Our code is available at https://github.com/iamjingli/LFM.

cs.CV↗

Universal Jamming Criticality and Self-Organizing Principles from Disorder to the Limit of Perfect Crystalline Order

While crystals are defined by periodic order, the nature of amorphous solids remains elusive due to their disordered, diverse, and nonequilibrium structures. Here, we focus on jammed elastic packings and systematically tune structure from crystalline to fully disordered to unveil the universal underlying characteristics. We demonstrate that their mechanical properties are universally governed by jamming criticality, featuring characteristic scaling behaviors near the jamming transition, excepting the singular close-packed point. This is facilitated by random nonaffine elasticity arising from contact-level disorder. Consequently, the jamming density can approach close packing, suggesting a fundamental decoupling between jamming criticality and the glass transition physics. Moreover, we uncover a universal coordination-number distribution and contact hyperuniformity in marginally jammed states, independent of particle-level structure. These findings suggest a general organizing mechanism for emergent rigidity in disordered solids, underscore the broad relevance of jamming physics, and complement principles of mechanical self-organization.

cond-mat.soft↗

DeepRT Engine: A Unified GPU-Parallel Ray-Tracing Framework with Hybrid SBR-IM Path Search for 6G Digital Twin Channel

Digital twin channel (DTC) aims to establish a real-time digital counterpart of physical wireless channels for reproducing and predicting site-specific propagation characteristics. As a high-precision channel computation method for realistic propagation scenarios, ray tracing (RT) serves as a key enabler for DTC construction. However, conventional RT suffers from high complexity under serial path-searching workflows. This letter proposes DeepRT Engine (DeepRT-E), a parallel RT acceleration architecture with a three-stage physically-inspired pipeline for real-time DTC construction. Firstly, DeepRT-E constructs a bounding volume hierarchy (BVH) to partition the scene and reduce redundant ray-surface intersections. Secondly, the shooting and bouncing rays (SBR) algorithm is executed through a ray-level parallel tracing framework to identify candidate surface sequences and prune the search space of the image method (IM). Finally, a parallel batched IM solver refines the retained candidates for accurate propagation-path recovery. Simulation results show that DeepRT-E reduces runtime by 96.3% and achieves a converged error of only 0.001 dB, outperforming Wireless InSite and Sionna in efficiency and accuracy.

eess.SP↗