SearcharxivSearch

arXiv subjects

Jibo Wei

Publications and source records attributed to Jibo Wei.

At least 19 recordsLinked to original sources

Personalized Digital Semantic Communication for Image Transmission with Vision-Language Models

Semantic communication (SC) enables bandwidth-efficient wireless image transmission, but most existing SC schemes are user-agnostic and ignore receiver-dependent semantics. To address this issue, we propose a personalized digital semantic communication (PDSC) framework that integrates a vision-language model (VLM)-based semantic encoder with a latent diffusion model (LDM)-based semantic decoder. Specifically, the semantic encoder extracts source-aware personalized semantic tokens from both the source image and the receiver's historical interactions. These tokens are vector-quantized into discrete semantic indices and further encoded into a compact fixed-length bitstream, enabling compatibility with digital transmission. At the receiver, the semantic decoder reconstructs a personalized image conditioned on the recovered semantic tokens. Furthermore, we formulate a capacity-constrained personalized semantic rate-distortion problem and introduce a semantic distortion metric that jointly characterizes source-semantic fidelity and user-preference alignment. Experiments show that PDSC achieves superior source-semantic consistency and personalization over state-of-the-art SC baselines, including CDDM and MoS, under bandwidth-limited wireless transmission.

eess.IV

Discovery of unobservable parameters via physical embedding

Recovering a source signal from indirect measurements often requires estimating latent parameters, such as wireless channel states or MRI coil sensitivities, that cannot be directly observed. Here, we introduce Physics-Embedded Inverse Learning (PEIL), in which a learned estimator predicts these parameters and a fixed, physics-based inverse operator uses them to reconstruct the signal, so that training requires only the source signal as supervision. In systems where multiple parameter combinations can reconstruct the signal equally well, the estimator exploits this freedom to coordinate parameters that compensate for residual modelling errors rather than match ground-truth parameters. In high-mobility wireless communications, PEIL discovers task-optimal configurations that outperform baselines given access to ground-truth parameters, enabling zero-shot generalisation and over 20-fold reduction in training data relative to supervised baselines. To test whether these properties extend across physical domains, we apply PEIL to parallel MRI, where it discovers physically interpretable coil sensitivity maps without calibration scans, yielding reconstructions grounded purely in acquired measurements. These results demonstrate that non-identifiability, conventionally a liability, becomes a resource when the learning objective targets reconstruction quality rather than parameter accuracy.

eess.SP

A Stage-Wise Learning Strategy with Fixed Anchors for Robust Speaker Verification

Learning robust speaker representations under noisy conditions presents significant challenges, which requires careful handling of both discriminative and noise-invariant properties. In this work, we proposed an anchor-based stage-wise learning strategy for robust speaker representation learning. Specifically, our approach begins by training a base model to establish discriminative speaker boundaries, and then extract anchor embeddings from this model as stable references. Finally, a copy of the base model is fine-tuned on noisy inputs, regularized by enforcing proximity to their corresponding fixed anchor embeddings to preserve speaker identity under distortion. Experimental results suggest that this strategy offers advantages over conventional joint optimization, particularly in maintaining discrimination while improving noise robustness. The proposed method demonstrates consistent improvements across various noise conditions, potentially due to its ability to handle boundary stabilization and variation suppression separately.

cs.SD

Noise-Conditioned Mixture-of-Experts Framework for Robust Speaker Verification

Robust speaker verification under noisy conditions remains an open challenge. Conventional deep learning methods learn a robust unified speaker representation space against diverse background noise and achieve significant improvement. In contrast, this paper presents a noise-conditioned mixture-ofexperts framework that decomposes the feature space into specialized noise-aware subspaces for speaker verification. Specifically, we propose a noise-conditioned expert routing mechanism, a universal model based expert specialization strategy, and an SNR-decaying curriculum learning protocol, collectively improving model robustness and generalization under diverse noise conditions. The proposed method can automatically route inputs to expert networks based on noise information derived from the inputs, where each expert targets distinct noise characteristics while preserving speaker identity information. Comprehensive experiments demonstrate consistent superiority over baselines

cs.SD

Theoretical Analysis of Deep Neural Networks in Physical Layer Communication

Recently, deep neural network (DNN)-based physical layer communication techniques have attracted considerable interest. Although their potential to enhance communication systems and superb performance have been validated by simulation experiments, little attention has been paid to the theoretical analysis. Specifically, most studies in the physical layer have tended to focus on the application of DNN models to wireless communication problems but not to theoretically understand how does a DNN work in a communication system. In this paper, we aim to quantitatively analyze why DNNs can achieve comparable performance in the physical layer comparing with traditional techniques, and also drive their cost in terms of computational complexity. To achieve this goal, we first analyze the encoding performance of a DNN-based transmitter and compare it to a traditional one. And then, we theoretically analyze the performance of DNN-based estimator and compare it with traditional estimators. Third, we investigate and validate how information is flown in a DNN-based communication system under the information theoretic concepts. Our analysis develops a concise way to open the "black box" of DNNs in physical layer communication, which can be applied to support the design of DNN-based intelligent communication techniques and help to provide explainable performance assessment.

eess.SP

Fine Timing and Frequency Synchronization for MIMO-OFDM: An Extreme Learning Approach

Multiple-input multiple-output orthogonal frequency-division multiplexing (MIMO-OFDM) is a key technology component in the evolution towards cognitive radio (CR) in next-generation communication in which the accuracy of timing and frequency synchronization significantly impacts the overall system performance. In this paper, we propose a novel scheme leveraging extreme learning machine (ELM) to achieve high-precision synchronization. Specifically, exploiting the preamble signals with synchronization offsets, two ELMs are incorporated into a traditional MIMO-OFDM system to estimate both the residual symbol timing offset (RSTO) and the residual carrier frequency offset (RCFO). The simulation results show that the performance of the proposed ELM-based synchronization scheme is superior to the traditional method under both additive white Gaussian noise (AWGN) and frequency selective fading channels. Furthermore, comparing with the existing machine learning based techniques, the proposed method shows outstanding performance without the requirement of perfect channel state information (CSI) and prohibitive computational complexity. Finally, the proposed method is robust in terms of the choice of channel parameters (e.g., number of paths) and also in terms of "generalization ability" from a machine learning standpoint.

eess.SP

Opening the Black Box of Deep Neural Networks in Physical Layer Communication

Deep Neural Network (DNN)-based physical layer techniques are attracting considerable interest due to their potential to enhance communication systems. However, most studies in the physical layer have tended to focus on the application of DNN models to wireless communication problems but not to theoretically understand how does a DNN work in a communication system. In this paper, we aim to quantitatively analyze why DNNs can achieve comparable performance in the physical layer comparing with traditional techniques and their cost in terms of computational complexity. We further investigate and also experimentally validate how information is flown in a DNN-based communication system under the information theoretic concepts.

eess.SP

Performance Analysis on Machine Learning-Based Channel Estimation

Recently, machine learning-based channel estimation has attracted much attention. The performance of machine learning-based estimation has been validated by simulation experiments. However, little attention has been paid to the theoretical performance analysis. In this paper, we investigate the mean square error (MSE) performance of machine learning-based estimation. Hypothesis testing is employed to analyze its MSE upper bound. Furthermore, we build a statistical model for hypothesis testing, which holds when the linear learning module with a low input dimension is used in machine learning-based channel estimation, and derive a clear analytical relation between the size of the training data and performance. Then, we simulate the machine learning-based channel estimation in orthogonal frequency division multiplexing (OFDM) systems to verify our analysis results. Finally, the design considerations for the situation where only limited training data is available are discussed. In this situation, our analysis results can be applied to assess the performance and support the design of machine learning-based channel estimation.

eess.SP

A Low Complexity Learning-based Channel Estimation for OFDM Systems with Online Training

In this paper, we devise a highly efficient machine learning-based channel estimation for orthogonal frequency division multiplexing (OFDM) systems, in which the training of the estimator is performed online. A simple learning module is employed for the proposed learning-based estimator. The training process is thus much faster and the required training data is reduced significantly. Besides, a training data construction approach utilizing least square (LS) estimation results is proposed so that the training data can be collected during the data transmission. The feasibility of this novel construction approach is verified by theoretical analysis and simulations. Based on this construction approach, two alternative training data generation schemes are proposed. One scheme transmits additional block pilot symbols to create training data, while the other scheme adopts a decision-directed method and does not require extra pilot overhead. Simulation results show the robustness of the proposed channel estimation method. Furthermore, the proposed method shows better adaptation to practical imperfections compared with the conventional minimum mean-square error (MMSE) channel estimation. It outperforms the existing machine learning-based channel estimation techniques under varying channel conditions.

eess.SP

Cooperative Multi-Agent Reinforcement Learning Based Distributed Dynamic Spectrum Access in Cognitive Radio Networks

With the development of the 5G and Internet of Things, amounts of wireless devices need to share the limited spectrum resources. Dynamic spectrum access (DSA) is a promising paradigm to remedy the problem of inefficient spectrum utilization brought upon by the historical command-and-control approach to spectrum allocation. In this paper, we investigate the distributed DSA problem for multi-user in a typical multi-channel cognitive radio network. The problem is formulated as a decentralized partially observable Markov decision process (Dec-POMDP), and we proposed a centralized off-line training and distributed on-line execution framework based on cooperative multi-agent reinforcement learning (MARL). We employ the deep recurrent Q-network (DRQN) to address the partial observability of the state for each cognitive user. The ultimate goal is to learn a cooperative strategy which maximizes the sum throughput of cognitive radio network in distributed fashion without coordination information exchange between cognitive users. Finally, we validate the proposed algorithm in various settings through extensive experiments. From the simulation results, we can observe that the proposed algorithm can converge fast and achieve almost the optimal performance.

cs.NI

Scalable Power Control/Beamforming in Heterogeneous Wireless Networks with Graph Neural Networks

Machine learning (ML) has been widely used for efficient resource allocation (RA) in wireless networks. Although superb performance is achieved on small and simple networks, most existing ML-based approaches are confronted with difficulties when heterogeneity occurs and network size expands. In this paper, specifically focusing on power control/beamforming (PC/BF) in heterogeneous device-to-device (D2D) networks, we propose a novel unsupervised learning-based framework named heterogeneous interference graph neural network (HIGNN) to handle these challenges. First, we characterize diversified link features and interference relations with heterogeneous graphs. Then, HIGNN is proposed to empower each link to obtain its individual transmission scheme after limited information exchange with neighboring links. It is noteworthy that HIGNN is scalable to wireless networks of growing sizes with robust performance after trained on small-sized networks. Numerical results show that compared with state-of-the-art benchmarks, HIGNN achieves much higher execution efficiency while providing strong performance.

cs.LG

Numerology Selection for OFDM Systems Based on Deep Neural Networks

In order to support diverse scenarios and deployments, the numerology of orthogonal frequency division multiplexing (OFDM) is defined for the parametrization of subcarrier spacing and cyclic prefix (CP). The time-frequency dispersion of mobile radio channels and the channel noise result in different performance deterioration in different numerologies. In this letter, we propose a deep neutral network (DNN) approach for numerology selection of OFDM systems. Considering the inter-symbol interference (ISI), inter-carrier interference (ICI) and noise level, the SNR loss is established as the objective to be minimized. We extract the power delay profile, mobile velocity and noise power as the input features to the DNN. The proposed DNN learns from the channel characteristics to obtain the optimal numerology selection. Simulation results show that the proposed DNN achieves better performance than the existing methods. The decision boundaries of different numerologies are also illustrated to show the application range according to the channel characteristics.

eess.SP

A Cyber Physical System Framework for UAV Communications

Diverse applications have witnessed the prevalence of unmanned aerial vehicles (UAVs) due to their agility and versatility. Compared with computation and control, the communication tends to be the bottleneck of the whole UAV system. Cyber physical system (CPS), which achieves the integration of the cyber and physical domains, can inspire us to deal with the communication problems through a cross-disciplinary method. To this end, we first expound the coupling effects of computation and control to communication. Then, we propose a novel CPS framework for UAV communications. By extending the dimension of communication decisions to computation and control, the framework can precisely orient and settle the communication issues. Further, a quantitative energy optimization model is established to guide the protocol and algorithm design for UAV communications. Case simulation results validate the CPS framework in terms of the energy consumption and communication delay.

eess.SP

Enhanced LMMSE Estimation Capable of Selecting Parameters

In the linear minimum mean square error (LMMSE) estimation for orthogonal frequency division multiplexing (OFDM) systems, the problem about the determination of the algorithm's parameters, especially those related with channel frequency response (CFR) correlation, has not been readily solved yet. Although many approaches have been proposed to determine the statistic parameters, it is hard to choose the best one within those approaches in the design phase, since every approach has its own most suitable application conditions and the real channel condition is unpredictable. In this paper, we propose an enhance LMMSE estimation capable of selecting parameters by itself. To this end, sampled noise MSE is first proposed to evaluate the practical performance of interpolation. Based on this evaluation index, a novel parameter comparison scheme is proposed to determine the parameters which can endow LMMSE estimation best performance within a parameter set. After that, the structure of the enhanced LMMSE is illustrated, and it is applied in OFDM systems. Besides, the issues about theoretical analysis on accuracy of the parameter comparison scheme, the parameter set design and algorithm complexity are explained in detail. At last, our analyses and performance of the proposed estimation method are demonstrated by simulation experiments.

eess.SP

Peak-to-Average Power Ratio Analysis for OFDM-Based Mixed-Numerology Transmissions

In this paper, the probability distribution of the peak to average power ratio (PAPR) is analyzed for the mixed numerologies transmission based on orthogonal frequency division multiplexing (OFDM). State of the art theoretical analysis implicitly assumes continuous and symmetric frequency spectrum of OFDM signals. Thus, it is difficult to be applied to the mixed-numerology system due to its complication. By comprehensively considering system parameters, including numerology, bandwidth and power level of each subband, we propose a generic analytical distribution function of PAPR for continuous-time signals based on level-crossing theory. The proposed approach can be applied to both conventional single numerology and mixed-numerology systems. In addition, it also ensures the validity for the noncontinuous-OFDM (NC-OFDM). Given the derived distribution expression, we further investigate the effect of power allocation between different numerologies on PAPR. Simulations are presented and show the good match of the proposed theoretical results.

eess.SP

PAPR Reduction Using Iterative Clipping/Filtering and ADMM Approaches for OFDM-Based Mixed-Numerology Systems

Mixed-numerology transmission is proposed to support a variety of communication scenarios with diverse requirements. However, as the orthogonal frequency division multiplexing (OFDM) remains as the basic waveform, the peak-to average power ratio (PAPR) problem is still cumbersome. In this paper, based on the iterative clipping and filtering (ICF) and optimization methods, we investigate the PAPR reduction in the mixed-numerology systems. We first illustrate that the direct extension of classical ICF brings about the accumulation of inter-numerology interference (INI) due to the repeated execution. By exploiting the clipping noise rather than the clipped signal, the noise-shaped ICF (NS-ICF) method is then proposed without increasing the INI. Next, we address the in-band distortion minimization problem subject to the PAPR constraint. By reformulation, the resulting model is separable in both the objective function and the constraints, and well suited for the alternating direction method of multipliers (ADMM) approach. The ADMM-based algorithms are then developed to split the original problem into several subproblems which can be easily solved with closed-form solutions. Furthermore, the applications of the proposed PAPR reduction methods combined with filtering and windowing techniques are also shown to be effective.

eess.SP

Survey on Unmanned Aerial Vehicle Networks: A Cyber Physical System Perspective

Unmanned aerial vehicle (UAV) networks are playing an important role in various areas due to their agility and versatility, which have attracted significant attention from both the academia and industry in recent years. As an integration of the embedded systems with communication devices, computation capabilities and control modules, the UAV network could build a closed loop from data perceiving, information exchanging, decision making to the final execution, which tightly integrates the cyber processes into the physical devices. Therefore, the UAV network could be considered as a cyber physical system (CPS). Revealing the coupling effects among the three interacted components in this CPS system, i.e., communication, computation and control, is envisioned as the key to properly utilize all the available resources and hence improve the performance of the UAV networks. In this paper, we present a comprehensive survey on the UAV networks from a CPS perspective. Firstly, we respectively research the basics and advances with respect to the three CPS components in the UAV networks. Then we look inside to investigate how these components contribute to the system performance by classifying the UAV networks into three hierarchies, i.e., the cell level, the system level, and the system of system level. Further, the coupling effects among these CPS components are explicitly illustrated, which could be enlightening to deal with the challenges in each individual aspect. New research directions and open issues are discussed at the end of this survey. With this intensive literature review, we try to provide a novel insight into the state-of-the-art in the UAV networks.

cs.NI

Deep Neural Network Aided Scenario Identification in Wireless Multi-path Fading Channels

This letter illustrates our preliminary works in deep nerual network (DNN) for wireless communication scenario identification in wireless multi-path fading channels. In this letter, six kinds of channel scenarios referring to COST 207 channel model have been performed. 100% identification accuracy has been observed given signal-to-noise (SNR) over 20dB whereas a 88.4% average accuracy has been obtained where SNR ranged from 0dB to 40dB. The proposed method has tested under fast time-varying conditions, which were similar with real world wireless multi-path fading channels, enabling it to work feasibly in practical scenario identification.

eess.SP