SearcharxivSearch

arXiv subjects

Zhen Gao

Publications and source records attributed to Zhen Gao.

At least 19 recordsLinked to original sources

From Semantic to Token Communication: The Next Paradigm for Large-Model-Driven 6G Intelligent Connectivity

The ambitious requirements of sixth-generation (6G) networks are driving communication systems from reliable bit delivery toward meaning-aware and task-oriented connectivity. Large models (LMs), with strong multimodal understanding and generation capabilities, have accelerated this shift and made semantic communication (SemCom) increasingly practical. Yet current LM-driven SemCom remains fragmented: semantic representations are typically tied to specific modalities, models, or tasks. While the bit provides a universal unit for digital transport, there is still no analogous unit for representing and processing semantics, which limits interoperability, theoretical unification, and scalable system design. We argue that tokens provide a natural candidate for this missing abstraction. Two trends support this: unified multimodal LMs now encode text, images, audio, video, and robot actions in one token space, while distributed LM inference already generates substantial token-level traffic through expert routing, cache transfer, and speculative decoding. Token communication (TokenCom) emerges by unifying these trends, using the LM's native processing unit as a communication abstraction above the bit level and enabling importance assignment, error handling, and resource allocation directly at token granularity. This survey traces the evolution from LM-driven SemCom to TokenCom. We review three major directions of LM-driven SemCom: source-centric semantic coding, channel semantics for physical-layer tasks, and collaborative edge-device intelligence. We then examine the token abstraction, the transmission techniques it requires, and two emerging paradigms, namely TokenCom for LM services and for embodied and agentic intelligence. Finally, we identify open challenges toward unified, scalable, and AI-native 6G communication systems.

eess.SP

Feature Reconfiguration With Visual Prior for Medical Lesion Segmentation

Lesion segmentation in medical images plays a critical role in clinical diagnosis and treatment planning. Despite significant advances, lesion segmentation remains challenging due to two major factors: (1) complex background interference; (2) diverse lesion morphology. Existing encoder-decoder based methods mainly focus on enhancing feature extraction or redesigning decoding strategies. However, they lack early prior guidance and feature reconfiguration during the encoding stage, limiting their effectiveness in handling these challenges. To address these limitations, we propose FreNet, a feature reconfiguration framework with visual priors, which performs pixel-level reconfiguration before encoding and feature-level reconfiguration during encoding for precise medical lesion segmentation. To suppress background responses, we propose an Implicit Prior Neural Network (IPNN), which models a continuous spatial field and leverages visual prior from SAM to reconfigure input image before encoding stage. To better handle diverse lesion morphology, we design a Dual-domain Feature Reconfiguration (DFR) module to progressively reconfigure backbone features during encoding stage. Within DFR, the Frequency Decoupling Module (FDM) decouples backbone features in frequency domain to enhance foreground-background discriminability, while the Spatial Localization Module (SLM) spatially relocates and improving spatial stability after frequency decoupling. Extensive experiments on 9 medical image segmentation benchmarks across three imaging modalities demonstrate that FreNet significantly outperforms state-of-the-art (SOTA) methods. On the challenging ETIS dataset, our method achieves Dice improvements of 5.0% over SOTA method and 7.2% over SAM.

cs.AI

TokenComSR: Task-Sensitivity-Guided Token Communication for Wireless Image Super-Resolution

For resource-constrained wireless edge devices over bandwidth-limited fading channels, wireless image transmission using traditional separate coding suffers from the cliff-effect collapse. Prevailing deep joint source-channel coding (JSCC) based on convolutional neural networks can mitigate this issue but usually fail to preserve patch-level structures, thereby preventing adaptive per-token power allocation and limiting token-domain compensation for super-resolution (SR). To address these challenges, we propose a token communication framework with SR (TokenComSR). Specifically, we conceive a task-sensitive power allocation (TSPA) module and a signal-to-noise ratio (SNR)-conditioned token refinement module (TRM). TSPA distills training estimates of task sensitivity into inference token power weights, while TRM estimates an SNR-conditioned residual to correct channel-induced distortion in the token domain before decoding. Building on TSPA and TRM, the proposed TokenComSR pairs a Swin Transformer-based token transceiver with a receiver-side SR module for resource-constrained wireless image transmission. Simulation results confirm the effectiveness of the proposed TSPA and TRM, demonstrating improvements over separate coding and JSCC-SR baselines in both reconstruction fidelity and perceptual quality.

cs.IT

Ada-TokenCom: Rate-Adaptive Token Communications via Large-Model-Driven Token Compression and Generation

Token Communications (TokenCom) has recently emerged as a new paradigm in which tokens serve as unified units for communication and computation, enabling efficient multimodal semantic and goal-oriented transmission. In this paper, we develop Ada-TokenCom, a rate-adaptive TokenCom framework based on large autoregressive models, which integrates next-token prediction with arithmetic coding to achieve ultra-low bitrate semantic communication at the token level. We propose a mixed reconstruction/generation scheme, where the transmitter encodes and transmits the highly informative tokens at the beginning of the token sequence leveraging a pre-trained autoregressive large model, while the receiver uses an identical model to predict the rest. Moreover, we design a Lyapunov-based algorithm to dynamically optimize both the source compression rate and the modulation and coding scheme, adapting to time-varying network conditions. Simulation results demonstrate that our proposed Ada-TokenCom framework outperforms both digital and deep joint source-channel coding-based semantic communication baselines.

cs.IT

Knowledge Distillation Driven Semantic NOMA with GAN Refinement for 6G Robotic Vehicle Networks

To achieve sustainable intelligent mobility, 6G-empowered robotic vehicles (RVs) require high-fidelity visual perception under stringent bandwidth and energy constraints. Semantic communication offers a spectral-efficient solution but suffers from severe interference in uplink non-orthogonal multiple access (NOMA) RV networks. To address this, we propose a knowledge distillation-driven and generative models-enhanced NOMA framework for robust and green RV communications, named KDG-SemNOMA. First, we develop a ConvNeXt-based deep joint source-channel coding (DeepJSCC) architecture with an enhanced attention feature (AF) module for dynamic channel adaptation. Second, to mitigate interference without inference overhead, an orthogonal transmission teacher model guides the NOMA student model via a two-stage knowledge distillation strategy. Finally, to address the over-smoothing artifacts of pixel-wise optimization, we introduce a channel-conditional GAN (cGAN). By explicitly taking the Stage-I initial reconstruction and channel states as conditional inputs, this module refines coarse outputs into high-fidelity images with realistic textures. Experiments on FFHQ-256 demonstrate that KDG-SemNOMA significantly outperforms state-of-the-art methods in both pixel-level accuracy and perceptual fidelity.

cs.IT

UW-OCDM for Low-Altitude UAV Communication and Cooperative Sensing

Integrated sensing and communications (ISAC) is a key enabler for uncrewed aerial vehicles (UAVs) in the low-altitude economy. This paper proposes an ISAC waveform that embeds a unique word (UW) into orthogonal chirp division multiplexing (OCDM), termed UW-OCDM, together with corresponding communication reception and cooperative sensing schemes for high-mobility UAV scenarios. For communication, the embedded UW enables timing synchronization and Doppler estimation and compensation without requiring a separate synchronization sequence. A sparse spatio-temporal channel estimation method exploits the common channel support across multiple receive antennas and consecutive UW observations to support reliable data demodulation. For sensing, the deterministic UW serves as a shared prior that allows distributed base stations to construct sensing dictionaries locally without exchanging random payload symbols in real time. A hierarchical multi-target detection and tracking algorithm integrates direct-path interference suppression, kinematic prediction, multi-candidate screening, off-grid refinement, residual verification, and successive interference cancellation for robust localization with reduced search complexity. Simulation results demonstrate reliable communication and localization in highly dynamic UAV scenarios, while the proposed framework retains low-complexity frequency-domain equalization and reduces transmit-reference sharing overhead and multi-static localization complexity.

eess.SP

Time-Reversal-Invariant Altermagnetic Acoustic Crystals

Altermagnets have emerged as a new class of magnetic materials that combine spin-split electronic bands with zero net magnetization. Extending this paradigm to classical-wave systems has, however, been fundamentally challenging because conventional realizations require broken time-reversal symmetry (TRS). Here, we overcome this limitation by introducing two pseudospin degrees of freedom and constructing a pseudo-time-reversal operator that faithfully reproduces the action of its physical counterpart while preserving actual TRS. Building on this framework, we theoretically propose and experimentally realize the first time-reversal-invariant altermagnetic acoustic crystal. Acoustic measurements directly reveal pseudospin-dependent band splitting--a defining hallmark of altermagnetism--under strictly TRS-preserving conditions. Moreover, the altermagnetic acoustic crystal exhibits sublattice-pseudospin locking, enabling flexible control over acoustic pseudospin splitting and filtering. Our work establishes acoustic crystals as a versatile platform for exploring altermagnetic physics and opens new avenues for spin-inspired wave manipulation in nonmagnetic devices.

cond-mat.mes-hall

Observation of Antichiral Hinge States in a Three-dimensional Gyromagnetic Photonic Crystal

Recent advances in topological physics have revealed a counterintuitive class of antichiral edge and surface states that propagate in the same direction along spatially separated parallel boundaries. To date, however, experimental realizations of antichiral states have been restricted to first-order topological phases, while their higher-order counterparts--antichiral hinge states--have remained experimentally elusive. Here, we report the first experimental observation of antichiral hinge states in a gyromagnetic photonic crystal that realizes a three-dimensional (3D) modified Haldane model with dimerized interlayer coupling. Through microwave near-field mapping, we directly resolve their defining signatures: nonreciprocal, co-propagating transport along four parallel hinges and characteristically tilted hinge-state dispersions. These results extend antichiral topology into the higher-order regime and provide a new platform for 3D nonreciprocal topological photonic devices.

physics.optics

Observation of biased random-flux-induced topological phase transition in gyromagnetic photonic crystals

The interplay between disorder and topological states has attracted growing interest. While previous studies have primarily addressed the effects of geometric or potential randomness, the exploration of topological phase transitions driven by random-flux remains experimentally elusive. Here, we report the first experimental realization of topological phase transitions driven by biased random-flux in gyromagnetic photonic crystals. By stochastically orienting the magnetization of constituent gyromagnetic rods, we implement a disordered Haldane model in which the sign of the next-nearest-neighbor hopping phases is randomly distributed. We demonstrate that the bulk band gap closes when the densities of positive and negative flux are balanced, i.e, restoring time-reversal symmetry in a statistical sense, and reopens when a net positive or negative flux is introduced. Microwave near-field measurements directly visualize the reversal of chiral edge states, confirming a transition between distinct topological phases. Our results establish a unique disorder-driven mechanism for realizing topological phase transitions and deepen our understanding of the interplay between disorder and topological phases in bosonic systems.

physics.optics

Multi-Domain Iterative Detection for Massive Connectivity in LEO Satellite Networks

Grant-Free (GF) random access is promising for low Earth orbit satellite Internet due to its reduced access latency. However, existing schemes suffer from poor performance in massive connectivity scenarios. To address this challenge, we firstly propose an iterative residual feedback multi-measurement vector approximate message passing algorithm. This algorithm leverages multi-domain synergistic sparsity in the spatial-frequency and angular-delay domains to alternately perform active user terminal detection (AUD) and channel estimation (CE). Additionally, a residual feedback mechanism is incorporated to suppress error accumulation, thereby enhancing AUD performance. Furthermore, conventional data detection (DD) methods significantly degrade when active user terminals are spatially close or outnumber the satellite's receive antennas, making the demodulation problem rank-deficient or underdetermined. To mitigate this, we design a data modulation scheme via joint spatial-frequency multi-domain spreading, which utilizes observations from both spatial and frequency domains to facilitate multi-domain DD. Simulation results demonstrate that the proposed scheme significantly outperforms existing GF methods in terms of AUD accuracy, CE precision, and bit error rate, especially under conditions of low effective pilot length and practical signal-to-noise ratios.

eess.SP

Fifth-Order Well-Balanced Path-Conservative A-WENO Scheme for the Ripa Model

In this work, we introduce a fifth-order well-balanced (WB) path-conservative A-WENO scheme with the central-upwind numerical fluxes (PCCU-5) for the Ripa model. The proposed scheme is capable of exactly preserving a variety of steady states, including still-water, moving-water, isobaric, and constant water height ones. This goal is achieved with the help of a flux globalization technique: The source terms are incorporated into the fluxes, resulting in a quasi-conservative system, for which central-upwind numerical fluxes are computed using the path-conservative integration. The proposed A-WENO scheme utilizes a WENO interpolation of the equilibrium variables rather than the conservative ones to ensure the WB property. In addition, we perform the WENO interpolation of the local characteristic equilibrium variables to mitigate numerical oscillations near discontinuities. We perform a series of numerical experiments, which demonstrate that the proposed fifth-order WB PCCU-5 scheme achieves high resolution and clearly outperforms its second-order counterpart. Our numerical results also demonstrate the importance of the local characteristic projection for significantly reducing (eliminating) numerical oscillations near discontinuities.

math.NA

AirTF: Over-the-Air Token Fusion for Task-Oriented Multi-Modal Token Communications

In the Internet of Vehicles (IoV), transmitting high-dimensional multi-modal sensory data to edge servers for time-sensitive tasks faces severe spectrum bottlenecks. To address this, we propose a foundation model-driven over-the-air token fusion (AirTF) framework for task-oriented multi-modal token communications. Unlike existing schemes for segmentation that rely on convolutional neural networks (CNNs) with limited local receptive fields, AirTF leverages vision transformer (ViT) encoders to extract globally contextualized semantic tokens from distributed heterogeneous sensors. By concurrently transmitting these spatially aligned tokens over a shared wireless channel, our framework exploits the superposition property of the multiple access channel to inherently fuse complementary multi-modal semantics (e.g., RGB and infrared) directly over the air. This mechanism significantly enhances spectral efficiency compared to orthogonal transmission. Furthermore, the integration of a pre-trained foundation model provides critical visual priors, effectively addressing the data-hungry nature of ViTs on limited, scenario-specific semantic segmentation datasets. Experiments demonstrate that AirTF consistently outperforms orthogonal transmission and CNN-based fusion baselines across AWGN and fading channels. Additional evaluations under a three-user setting, residual synchronization errors, and imperfect channel state information estimation further confirm its robustness. The source code will be made publicly available upon acceptance.

eess.IV

Observation of fractality-induced topology in photonic crystals

Fractal topology--achieved by integrating nontrivial topology into fractal geometries with self-similarity and non-integer dimensions--has opened new avenues for exploring topological phases of matter. Recent theoretical advances revealed a counterintuitive fractal topology: fractality itself can induce nontrivial topology in an otherwise trivial system. Here, we report the first experimental observation of fractality-induced topology in a tight-binding-like photonic crystal, without relying on traditional driving mechanisms such as magnetic fields, staggered hopping, or spin-orbit coupling. We demonstrate that fractality alone is sufficient to lift the degeneracy of Kagome lattice band structure and induce topological corner states within the bandgap of the resulting fractal Kagome photonic crystal, which is a photonic higher-order topological insulator. This work experimentally reveals a novel mechanism for realizing nontrivial topological states, expanding both the fundamental frontier and potential application of topological physics.

physics.optics

LGVSC: A Large-Model-Driven Generative Video Semantic Communication Framework

Driven by the massive video transmission requirements in the Internet of Everything, semantic communication holds great promise for striking a balance between transmission efficiency and quality. This paper introduces a large-model-driven generative video semantic communication (LGVSC) framework, enabling efficient video semantic transmission under extremely low bandwidth conditions. First, by decoupling the encoder and decoder as well as exposing explicit intermediate semantic representations, LGVSC maintains interpretability, avoiding the black-box behavior commonly observed in end-to-end systems. Next, we introduce a new metric, i.e., the probability-based semantic similarity score (PSSS), which quantifies semantic similarity for complex modalities within a continuous range, allowing for more precise evaluation of semantic content. Building on PSSS, we propose a semantic-guided keyframe extraction module driven by a multimodal large model. This module can enhance fine-grained semantic consistency during keyframe selection at the transmitter, optimizing transmission bandwidth without compromising semantic fidelity. Additionally, we design a generative large-model-driven dynamic semantic-adaptive decoder at the receiver, which can adapt to videos of arbitrary lengths. Simulation results demonstrate that LGVSC significantly outperforms traditional schemes, achieving a channel bandwidth ratio on the order of $10^{-4}$ to $10^{-3}$, while maintaining strong zero-shot generalization across downstream tasks.

eess.SP

Risk Assessment of Autonomous Driving: Integrating Technical Failures, Ethical Dilemmas, and Policy Frameworks

Autonomous driving technology has the potential to reduce the large number of road traffic accidents caused by human error each year, but it also brings new types of risks that need to be evaluated from the aspects of technology, ethics and regulations. Based on public crash data from the National Highway Traffic Safety Administration (NHTSA), disengagement reports from the California Department of Motor Vehicles (DMV), the MIT Moral Machines dataset, and a comparative regulatory analysis of five jurisdictions, we have found that the main types of technical failure modes are perception and classification errors. These account for a relatively large proportion of the reported accidents, and it can be concluded that there are different ethical frameworks for autonomous vehicle decision-making, and inconsistent regulations in different areas increase the uncertainty of widespread application. Generally speaking, the problems of technology, ethics and regulation are closely related and need to be solved together. Therefore, this paper recommends a more adaptive and cooperative governance approach that combines engineering standards, ethical discussion, and institutional supervision.

cs.AI

ReFLEX: Length-Generalizable CSI Denoising for MIMO-OFDM via Relative-Frequency Bias

This letter studies CSI denoising for MIMO--OFDM with variable NR resource block (RB) allocations. ReFLEX is a length-generalizable Transformer whose frequency attention uses a relative-frequency position bias (RFPB) generated from subcarrier offsets. A single checkpoint handles unseen RB lengths and can be applied to sparse DM-RS observations in the tested RB5/RB10 PUSCH setup without retraining. In a 3GPP~TR~38.901 UMa NLOS channel, ReFLEX achieves about $-9.6$~dB NMSE on unseen RB lengths. In NR PUSCH/UL-SCH simulations, ReFLEX denoising followed by time-frequency interpolation reduces the 10\% BLER threshold by about 2--3~dB.

eess.SP

Orbital Altermagnetic Photonic Crystal

Altermagnetism features momentum-dependent spin splitting without net magnetization, extending spintronics beyond conventional ferromagnetism and antiferromagnetism. However, the photonic realization of altermagnetism has remained a formidable challenge due to the fundamental differences between fermionic electrons and bosonic photons. Here, we report the first experimental realization of an orbital altermagnetic photonic crystal, based on an antiunitary $C_{4z}\mathcal{T}$ symmetry enforced correspondence between a local $p$-orbital $\sigma/\pi$ doublet and crystal momentum. We experimentally demonstrate that the resulting system exhibits momentum-dependent spin splitting with alternating pseudospin polarization and a $d_{xy}$-wave form factor, as confirmed by measured band structures and iso-frequency contours. Moreover, we show that the orbital altermagnetic photonic crystal supports unique pseudospin-selective transport of electromagnetic waves, including photonic pseudospin splitting and pseudospin filtering. Our results extend the field of alternagnetism to photonic systems, opening a new avenue for designing spinphotonic devices.

physics.optics

Tri-Domain Multiuser MIMO Precoding Optimization and Channel Estimation with Spatial-EM Reconfigurable Antenna

In this paper, we propose a tri-domain reconfigurable multiuser multiple-input multiple-output (MIMO) communication system that integrates the electromagnetic (EM) reconfigurable antenna (EMRA) with the spatially movable antenna (SMA), termed the spatial-EM reconfigurable antenna (SEMRA). The proposed system offers EM, spatial, and digital domain degrees of freedom (DoFs) for joint channel reconfiguration, yet introduces new challenges in channel estimation (CE) and precoding optimization. Specifically, for multiuser orthogonal frequency division multiplexing (OFDM) downlink, the precoding design is formulated as a tri-domain optimization problem over antenna positions, EM-domain radiation-pattern weights, and digital precoders. We first develop a zero-forcing (ZF)-based baseline algorithm to decouple the design of spatial reconfiguration, and then propose a weighted minimum mean square error (WMMSE)-based tri-domain joint optimization algorithm for further improving the spectral efficiency (SE). Furthermore, we propose a low-overhead movement-aided channel estimation scheme in which coordinated antenna repositioning across pilot slots synthesizes a denser virtual array, enabling more accurate angle-of-departure (AoD) estimation and EM-domain channel state information (eCSI) reconstruction under the same per-user pilot overhead as the EMRA baseline. The resulting parametric representation enables eCSI assembly at desired antenna positions without additional pilots. Simulation results show that the proposed CE scheme improves eCSI estimation accuracy and the proposed SEMRA achieves higher SE than the EMRA baseline under the same pilot overhead.

eess.SP