SearcharxivSearch

arXiv subjects

Haoshuo Zhang

Publications and source records attributed to Haoshuo Zhang.

4 recordsLinked to original sources

Lightweight Generative Image Semantic Communication over Packet Erasure Channels

This paper addresses packet loss in semantic communication caused by network congestion or channel fluctuations. We propose LGSemCom, a lightweight generative packet-level joint source-channel coding (JSCC) framework for efficient and robust image transmission over packet erasure channels. Unlike conventional distortion-oriented recovery methods that yield blurry averages over erased regions, LGSemCom reformulates image recovery under packet erasures as a conditional generative task. By leveraging adversarial learning, the proposed decoder synthesizes plausible details without increasing complexity during inference. A key innovation of our framework is an erasure-aware weighting (EAW) strategy, which prioritizes generation in erased regions while preserving pixel fidelity in correctly received areas. To ensure computational efficiency, LGSemCom employs a fully convolutional codec based on efficient long-range attention blocks (ELABs) that capture global semantic dependencies with low complexity. Extensive experiments show that LGSemCom achieves superior perceptual reconstruction quality compared with existing benchmarks under severe packet loss, while achieving an order of magnitude faster inference. These attributes make LGSemCom highly suitable for latency-sensitive applications on resource-constrained edge devices.

eess.SP

BitSemCom: A Bit-Level Semantic Communication Framework with Learnable Probabilistic Mapping

Most existing semantic communication systems based on joint source-channel coding (JSCC) employ analog modulation and are thus inherently incompatible with modern digital communication systems and impose stringent hardware design challenges. Although several digital transmission approaches have been proposed to address this issue, they often suffer from high sensitivity to bit errors, limited adaptability to varying source distributions, or re-training overhead under different modulation schemes. This letter proposes BitSemCom, a novel end-to-end bit-level JSCC framework that is robust to channel noise and modulation-agnostic. The core component is a learnable bit mapper that establishes a probabilistic mapping between continuous semantic features and discrete bit sequences. By leveraging a sampling-based bit generation method based on the Gumbel-Softmax trick, the framework enables differentiable bit-level optimization while maintaining robustness to channel errors. Simulation results on image transmission demonstrate that BitSemCom achieves consistent peak signal-to-noise ratio (PSNR) gains of 2-3 dB over codebook-based digital semantic transmission methods and competitive performance with stronger robustness compared to separate source-channel coding (SSCC) benchmarks. Ablation studies further validate the effectiveness of the learnable bit mapper.

eess.IV

ProMSC-MIS: Prompt-based Multimodal Semantic Communication for Multi-Spectral Image Segmentation

Multimodal semantic communication has great potential to enhance downstream task performance by integrating complementary information across modalities. This paper introduces ProMSC-MIS, a novel Prompt-based Multimodal Semantic Communication framework for Multi-Spectral Image Segmentation. It enables efficient task-oriented transmission of spatially aligned RGB and thermal images over band-limited channels. Our framework has two main design novelties. First, by leveraging prompt learning and contrastive learning, unimodal semantic encoders are pre-trained to learn diverse and complementary semantic representations by using features from one modality as prompts for another. Second, a semantic fusion module that combines cross-attention mechanism and squeeze-and-excitation (SE) networks is designed to effectively fuse cross-modal features. Experimental results demonstrate that ProMSC-MIS substantially outperforms conventional image transmission combined with state-of-the-art segmentation methods. Notably, it reduces the required channel bandwidth by 50%--70% at the same segmentation performance, while also decreasing the storage overhead and computational complexity by 26% and 37%, respectively. Ablation studies also validate the effectiveness of the proposed pre-training and semantic fusion strategies. Our scheme is highly suitable for applications such as autonomous driving and nighttime surveillance.

cs.MM

Prompt-based Multimodal Semantic Communication for Multi-spectral Image Segmentation

Multimodal semantic communication has gained widespread attention due to its ability to enhance downstream task performance. A key challenge in such systems is the effective fusion of features from different modalities, which requires the extraction of rich and diverse semantic representations from each modality. To this end, we propose ProMSC-MIS, a Prompt-based Multimodal Semantic Communication system for Multi-spectral Image Segmentation. Specifically, we propose a pre-training algorithm where features from one modality serve as prompts for another, guiding unimodal semantic encoders to learn diverse and complementary semantic representations. We further introduce a semantic fusion module that combines cross-attention mechanisms and squeeze-and-excitation (SE) networks to effectively fuse cross-modal features. Simulation results show that ProMSC-MIS significantly outperforms benchmark methods across various channel-source compression levels, while maintaining low computational complexity and storage overhead. Our scheme has great potential for applications such as autonomous driving and nighttime surveillance.

eess.IV