SearcharxivSearch

arXiv subjects

Yin Xu

Publications and source records attributed to Yin Xu.

At least 19 recordsLinked to original sources

Tensor Decomposition Based Mixed-Field Sensing for XL-MIMO AFDM Systems

Integrated sensing and communications enabled by extremely large-scale MIMO (XL-MIMO) and affine frequency division multiplexing (AFDM) is a highly promising paradigm for vehicular networks. However, the near-field spherical wavefront distortions induce severe non-linear parameter coupling, while the highly dynamic scattering environments exacerbate mismatch errors. To address these critical challenges, this paper proposes a novel tensor-based sensing scheme for XL-MIMO AFDM systems. First, the received signals are reformulated into a tensor, followed by an efficient decomposition approach that exploits the inherent Vandermonde structure of the factor matrices. This allows parameters to be directly estimated from the decomposed matrices, effectively avoiding inter-parameter coupling. Subsequently, a symmetric decoupling and real-domain manifold optimization algorithm is proposed for angle of arrival estimation, circumventing the high-dimensional searches typically induced by near-field effects. Furthermore, a baseband reconstruction and analytical gradient-based algorithm is developed to perform delay-Doppler estimation in the continuous parameter domain, fundamentally eradicating the grid-mismatch errors inherent in high-mobility scenarios. With these decoupled factors, the remaining unknown angle of departure can be readily extracted. Extensive simulation results demonstrate that the proposed scheme achieves orders-of-magnitude improvements in delay-Doppler accuracy and eliminates the error floors in angular estimation that severely bottleneck state-of-the-art baselines.

eess.SP

From Semantic to Token Communication: The Next Paradigm for Large-Model-Driven 6G Intelligent Connectivity

The ambitious requirements of sixth-generation (6G) networks are driving communication systems from reliable bit delivery toward meaning-aware and task-oriented connectivity. Large models (LMs), with strong multimodal understanding and generation capabilities, have accelerated this shift and made semantic communication (SemCom) increasingly practical. Yet current LM-driven SemCom remains fragmented: semantic representations are typically tied to specific modalities, models, or tasks. While the bit provides a universal unit for digital transport, there is still no analogous unit for representing and processing semantics, which limits interoperability, theoretical unification, and scalable system design. We argue that tokens provide a natural candidate for this missing abstraction. Two trends support this: unified multimodal LMs now encode text, images, audio, video, and robot actions in one token space, while distributed LM inference already generates substantial token-level traffic through expert routing, cache transfer, and speculative decoding. Token communication (TokenCom) emerges by unifying these trends, using the LM's native processing unit as a communication abstraction above the bit level and enabling importance assignment, error handling, and resource allocation directly at token granularity. This survey traces the evolution from LM-driven SemCom to TokenCom. We review three major directions of LM-driven SemCom: source-centric semantic coding, channel semantics for physical-layer tasks, and collaborative edge-device intelligence. We then examine the token abstraction, the transmission techniques it requires, and two emerging paradigms, namely TokenCom for LM services and for embodied and agentic intelligence. Finally, we identify open challenges toward unified, scalable, and AI-native 6G communication systems.

eess.SP

Agent-Native Task-Oriented Communication with Joint Token Compression Coding and Modulation

As large foundation models empower agents to become pervasive across industries and emerge as central actors in intelligent systems, a fundamental rethinking of communication paradigms toward AI-native, agent-centric designs in the post-Shannon era becomes inevitable. One essential shift is that tokens, which are the minimal semantic units natively processed by large language models (LLMs), should replace bits as the fundamental unit of communication. However, existing works in the LLMs field assume lossless token transmission over high-speed wired links and largely neglect the air-interface overhead and channel distortions inherent in wireless environments, lacking a native design for wireless token communication systems. To bridge this gap, we propose an innovative design for a token transmitter-receiver architecture that facilitates task-oriented token transmission. Specifically, we propose JTCM (Joint Token Coding and Modulation), an AI-native semantic communication framework that jointly optimizes token representation, channel coding, and modulation to maximize downstream task performance directly. Correspondingly, we propose a two-stage training scheme. In the first stage, the token encoder-decoder pair is pre-trained to enable semantic-preserving compression and reconstruction. In the second stage, it is fine-tuned end-to-end with a multi-modal foundation model under specific downstream tasks to achieve task-aware optimization. Extensive experiments demonstrate that JTCM significantly reduces transmission overhead while enhancing task accuracy and robustness compared to state-of-the-art baselines in bandwidth- and SNR-constrained wireless channels.

eess.SP

Learned Blockwise Port Activation for Real Time Beamforming in Fluid Antenna Arrays

Fluid antenna arrays (FAAs), support multiuser downlink transmission by activating a subset of reconfigurable ports. The activation mask jointly determines the effective channel and the sparse radiating aperture, which requires a balance among sum rate, sidelobe suppression, hardware constraints, and online complexity. Channel driven selection can cluster active ports and increase sidelobes, whereas sidelobe oriented synthesis is typically channel independent and can sacrifice sum rate. This paper proposes learned blockwise port activation (L-BPA), for real time sidelobe aware FAA downlink beamforming. L-BPA activates a fixed number of ports in each aperture block, which supports grouped switching hardware and limits port clustering. A lightweight convolutional network scores ports using multiuser channel features, port coordinates, and user power statistics. Training combines blockwise straight through masks with a differentiable peak sidelobe level (PSLL), surrogate. During inference, learned scores are combined with multiscale geometric repulsion, followed by regularized zero forcing precoding over the reduced effective channel. L-BPA reduces the average PSLL by 3.26 dB relative to uniform sparse activation while achieving a slightly higher sum rate. It also reduces the PSLL by 8.13 dB and 10.10 dB relative to greedy and gain based selection, respectively, without iterative online search.

cs.IT

Spatial Semantic Communication: When Semantic Transmission Meets Index Modulation

Current digital semantic communication systems have primarily focused on maintaining compatibility with conventional constellation-based modulation. In contrast, index modulation (IM) represents a more spectrally and energy-efficient alternative by exploiting additional dimensions for information conveyance. Recognizing this potential, this paper bridges the gap between IM and semantic communications by proposing a novel spatial semantic communication (SSC) system leveraging cutting-edge fluid antenna-IM (FA-IM) technology. Compatible with existing joint source-channel coding (JSCC) architectures, the proposed SSC system employs the residual quantization (RQ) approach to discretize analog semantic features for subsequent digital IM transmission. Notably, the proposed SSC system synergizes RQ and IM via a semantic-aware stream splitting scheme, which ensures that critical semantic information undergoes less severe channel fading, thereby further optimizing semantic transmission performance. Simulation results validate that the proposed SSC system effectively integrates the high fidelity of RQ, the reliability of semantic-aware splitting, and the spatial efficiency of FA-IM, thereby providing a robust solution for future digital semantic transmission. The open source code is available at: https://github.com/gxh1106/SSC.

cs.IT

MCRB and MSE Analysis for Parameter Estimation in AFDM-ISAC Systems

Affine frequency division multiplexing (AFDM) is a promising waveform for integrated sensing and communication (ISAC). In AFDM systems, the complex gains, delays, and Doppler shifts are commonly estimated from the AFDM symbols carrying pilots and data simultaneously. In practice, however, the unknown data symbols and data-pilot coupling interference may render the estimator mismatched to the true signal model. In this paper, we systematically characterize the parameter-estimation performance of AFDM-ISAC systems under practical model misspecification. The main contributions are threefold. First, we extend the Cram\'er-Rao bound (CRB) for a general observation model that treats the data symbols as unknown, which generalizes existing AFDM CRB analyses and serves as the matched benchmark for the subsequent analysis. Second, we identify two practical sources of misspecification, namely a covariance mismatch caused by insufficient pilot-data isolation and a combined covariance-and-mean mismatch caused by sequential single-target estimation, and derive the corresponding misspecified CRB (MCRB). Third, we characterize the pseudotrue parameters under different levels of prior knowledge, analyze the resulting estimation bias, and establish a lower bound (LB) on the mean square error (MSE). Simulation results validate the derived bounds and show that, under model misspecification, the CRB is overly optimistic while the MCRB and LB faithfully characterize the achievable accuracy. The comparison further reveals how these bounds vary with the pilot length and pilot power, providing useful guidance for pilot configuration.

eess.SP

Low-Complexity Hybrid Precoding for Cell-Free Massive MU-MIMO ISAC Systems

Integrated sensing and communication (ISAC) in cell-free (CF) massive multi-user multiple-input multiple-output (MU-MIMO) system is a promising architecture for high-rate communications and high-accuracy multi-target sensing. However, centralized coordination among distributed access points (APs) incurs substantial fronthaul overhead and computation complexity. This paper proposes a low-complexity hybrid precoding framework for CF massive MU-MIMO ISAC systems with partially-connected architectures at the APs. By applying hybrid architecture at the APs, the proposed framework converts the original high-dimensional channel information into a low-dimensional effective channel, enabling digital precoding over the compressed channel domain and thereby substantially reducing both fronthaul overhead and baseband computational complexity. We formulate the joint hybrid precoding design as an ergodic sum-rate (ESR) maximization problem with position error bound (PEB) constraints to ensure multi-target sensing accuracy. An efficient alternating optimization (AO)-based solver is then developed, where the PEB constraint is reformulated into tractable convex constraints, while the digital-domain optimization is carried out over the reduced-dimensional effective channel and the analog precoding is refined on the constant-modulus manifold. For dynamic user topology, we further propose multi-branch (MB) rate-splitting (RS) minimum mean-square-error Tomlinson-Harashima precoding (MMSE-THP) update algorithm that combines multi-branch ordering with recursive MMSE-THP matrix updates, enabling common and private digital precodings to be refreshed without repeated full matrix recomputation. Simulation results demonstrate that the proposed scheme achieves high ESR and accurate multi-target sensing while reducing computational complexity by 87.02\% compared with conventional baselines.

cs.IT

LGVSC: A Large-Model-Driven Generative Video Semantic Communication Framework

Driven by the massive video transmission requirements in the Internet of Everything, semantic communication holds great promise for striking a balance between transmission efficiency and quality. This paper introduces a large-model-driven generative video semantic communication (LGVSC) framework, enabling efficient video semantic transmission under extremely low bandwidth conditions. First, by decoupling the encoder and decoder as well as exposing explicit intermediate semantic representations, LGVSC maintains interpretability, avoiding the black-box behavior commonly observed in end-to-end systems. Next, we introduce a new metric, i.e., the probability-based semantic similarity score (PSSS), which quantifies semantic similarity for complex modalities within a continuous range, allowing for more precise evaluation of semantic content. Building on PSSS, we propose a semantic-guided keyframe extraction module driven by a multimodal large model. This module can enhance fine-grained semantic consistency during keyframe selection at the transmitter, optimizing transmission bandwidth without compromising semantic fidelity. Additionally, we design a generative large-model-driven dynamic semantic-adaptive decoder at the receiver, which can adapt to videos of arbitrary lengths. Simulation results demonstrate that LGVSC significantly outperforms traditional schemes, achieving a channel bandwidth ratio on the order of $10^{-4}$ to $10^{-3}$, while maintaining strong zero-shot generalization across downstream tasks.

eess.SP

Towards Standardizing Affine Frequency Division Multiplexing (AFDM) for Future Wireless Networks

Affine frequency division multiplexing~(AFDM) has emerged as a compelling waveform candidate for future wireless networks, owing to its strong resilience to doubly selective channels and its ability to enable the seamless integration of communication and sensing functionalities. Against this context, this article provides a systematic study of AFDM from a standardization perspective. We first introduce the principles of AFDM and discuss the major considerations involved in waveform standardization. We then examine the backwards compatibility of AFDM with 4G/5G multi-numerology frameworks and their anticipated evolution, frequency-modulated continuous-wave (FMCW) radar waveforms, and long-range (LoRa) modulation, demonstrating that AFDM can be incorporated into legacy processing chains with limited modification. Key standardization-critical capabilities are further discussed, including multiple-antenna and multi-user support, and peak-to-average power ratio (PAPR). Finally, we investigate the potential of AFDM in several emerging scenarios, including non-terrestrial networks~(NTN), integrated sensing and communications (ISAC), vehicle-to-everything (V2X), and underwater acoustic (UWA) communications, whereby severe delay-Doppler dispersion places stringent demands on waveform robustness. Through these explorations, it is shown that that AFDM represents a timely and compelling technology for future wireless networks.

eess.SP

SkySense: A Semi-Supervised Generative Framework for UAV Localization in ISAC Networks

Extreme data scarcity and inherent multipath spatial ambiguity severely limit existing deep learning-based channel state information (CSI) fingerprinting localization schemes for target unmanned aerial vehicles (UAVs). To overcome these challenges, we propose an end-to-end semi-supervised generative localization framework. First, by exploiting the temporal correlations inherent in continuous flight trajectories, a self-supervised encoder extracts robust spatial features from massive unlabeled CSI sequences to establish structured latent representations. Following this, we utilize a consistency model, a powerful derivative of diffusion architectures, as the core generative backbone to map the learned latent space to physical coordinates, jointly fine-tuning the pre-trained encoder with a strictly limited set of labeled CSI. This consistency formulation models the conditional distribution to resolve the mean collapse problem of discriminative models, while compressing the inference trajectory to 1-2 steps to avoid the latency bottleneck of traditional diffusion models. Furthermore, a lightweight distributed fusion mechanism is designed to aggregate spatial predictions across multiple base stations (BS) from a multi-view geometry perspective. Comprehensive evaluations on a real-world measurement dataset demonstrate that our framework achieves low latency and suppresses the mean localization error to 9.77 cm under a 3-BS fusion setup with only a 1\% label fraction, significantly outperforming existing fully supervised and semi-supervised discriminative baselines.

eess.SP

Scalable GNN-Based Power Allocation for Rate-Splitting Cell-Free Massive MIMO Systems

Cell-free massive multiple-input multiple-output (CF-mMIMO) systems provide enhanced coverage and capacity for next-generation wireless networks. However, CF-mMIMO systems face significant challenges in downlink power allocation (PA) due to imperfect channel state information (CSI), severe multi-user interference (MUI), and high computational complexity. To address these issues, rate-splitting multiple access (RSMA) is adopted as a robust interference management strategy. Accordingly, this paper proposes an unsupervised and scalable graph neural network (GNN) framework for PA in rate-splitting CF-mMIMO (RS-CF-mMIMO) systems, relying exclusively on large-scale fading (LSF) coefficients without instantaneous CSI. To resolve the dimensionality mismatch in dynamic networks, we introduce a slice-based adaptive layer that projects variable-dimension features into a fixed latent space. This mechanism enables a unified model to generalize across diverse topologies without retraining. Within this architecture, the sum spectral efficiency (SE) is maximized under per-AP power constraints, assuming maximum-ratio precoding for common streams and regularized zero-forcing precoding for private streams. We also derive a weighted minimum mean-square error-alternating direction method of multipliers (WMMSE-ADMM) algorithm as a performance upper bound. Extensive simulations verify that the proposed GNN framework achieves near-optimal SE and outperforms unsupervised deep neural networks (DNNs) across diverse system sizes and pilot assignment schemes. Furthermore, the scalable variant maintains robust performance while reducing the trainable parameter count by over 57% relative to DNNs and decreasing inference latency by up to three orders of magnitude compared with WMMSE-ADMM.

eess.SP

A NISQ-Aware Hybrid Quantum-Classical Framework for Scalable Combinatorial Optimization

Scalable combinatorial optimization under resource-constrained quantum hardware remains a fundamental challenge in the Noisy Intermediate-Scale Quantum (NISQ) era, due to the mismatch between exponentially growing solution spaces and limited quantum computational capacity. In this work, we propose a NISQ-aware hybrid quantum-classical optimization framework that reformulates large-scale combinatorial optimization as a resource-bounded distribution evolution process. Instead of directly optimizing individual solutions, the proposed framework operates on a probabilistic representation of the solution space, enabling efficient exploration under hardware constraints. Specifically, large problem instances are decomposed into qubit-compatible subproblems via clustering-based decomposition, ensuring resource-bounded optimization. Within each subproblem, a quantum genetic algorithm evolves the solution distribution, while periodically embedded amplitude amplification acts as a controlled quantum enhancement mechanism that accelerates convergence without increasing circuit depth. A classical refinement stage ensures global solution consistency. Extensive experiments on benchmark and synthetic datasets demonstrate that the proposed framework consistently outperforms classical and quantum-inspired baselines, with performance gains that become more pronounced as problem scale increases. This scale-dependent behavior indicates that scalability is achieved through structured decomposition rather than increased quantum complexity. Noise simulations further confirm robustness under realistic NISQ conditions, and ablation studies validate that both quantum evolutionary search and amplitude amplification contribute significantly to performance improvements.

quant-ph

MACReD: A Multi-Agent Collaborative Reasoning Framework for Reaction Diagram Parsing

Parsing chemical reaction diagrams from scientific literature is challenging due to heterogeneous layouts, intertwined visual elements, and the difficulty of integrating recognition and reasoning. Existing vision-language models advance multimodal understanding but still fail on complex diagrams, struggling to maintain spatial coherence and to integrate multidimensional information during reasoning. To address these issues, we propose MACReD, a hierarchical multi-agent framework that coordinates specialized agents for molecular perception, arrow understanding, text extraction, and reaction reconstruction within a unified VLM-guided architecture. The planning and perception layers use flexible, fine-grained detection to handle visual complexity, while the reasoning layer uses a multigraph fusion mechanism to integrate heterogeneous cues and enforce chemically consistent global reasoning. Experiments on the RxnScribe benchmark show that MACReD achieves state-of-the-art performance, with F1 scores of 75.2% and 84.6% under hard and soft match criteria, outperforming the RxnScribe baseline, which obtains 69.1% and 80.0%, respectively. These results demonstrate the robustness of MACReD across diverse diagram layouts, including multi-step and tree-structured reactions.

cs.AI

Waveform Design for 6G ISAC Systems Under Full-Duplex Residual Self-Interference

In this paper, the waveform design for 6G integrated sensing and communication (ISAC) systems is investigated, with a particular focus on the practical limitations imposed by imperfect full-duplex radios. Under such imperfections, continuous communication waveforms, such as OFDM, suffer from severe full-duplex residual self-interference (RSI) for radar sensing, which significantly restricts the long-range sensing capabilities required by emerging low-altitude wireless networks (LAWN). To address this challenge, we propose a novel time-division ISAC waveform that integrates a specially developed dual-power phase-coded pulse for sensing into the communication frame under full-duplex RSI. Specifically, the dual-power sensing pulse consists of a high-power sequence followed by a low-power sequence, effectively exploiting imperfect full-duplex operations to achieve reliable long-range sensing while eliminating the detection blind range inherent to conventional half-duplex pulse radars. Furthermore, a complementary and inverse-phase sequence group is designed to ensure perfect autocorrelation and robust cross-correlation sidelobe suppression, so as to enhance multi-target detection capability. As for sensing signal processing, a parameterized mismatched filter is developed and optimized to maximize the detection performance, tailored to the proposed pulse structure. In addition, we design a hierarchical one-dimensional CFAR-CA detector that can exploit the perfect range-domain autocorrelation characteristics of the proposed waveform to further improve the detection performance. Extensive simulations demonstrate that the proposed design significantly improves the maximum detection range and multi-target detection capability compared to existing OFDM and LFM pulse baselines, while effectively covering the blind range for targets with small RCS.

eess.SP

Quantum-Enhanced Recurrent Neural Networks via Variational Quantum Gating for Battery State of Health Prediction

Accurate state-of-health (SOH) estimation for lithium-ion batteries remains a challenging problem due to complex electrochemical degradation mechanisms and long-range temporal dependencies. In this work, we propose a quantum-enhanced recurrent framework, termed QLSTM, in which variational quantum circuits are directly embedded into the gating mechanisms of long short-term memory networks. By replacing classical affine transformations with parameterized unitary operations, the proposed model introduces structured nonlinear transformations into the recurrent state-transition process. Extensive experiments on multiple benchmark battery datasets demonstrate that QLSTM consistently outperforms classical sequence models in both predictive accuracy and robustness, achieving significant reductions in mean absolute error (MAE), with improvements on the order of 20% compared with classical LSTM baselines. Ablation studies further confirm that these improvements arise primarily from quantum-enhanced gating rather than input-level transformations. Additional analyses on qubit scaling and noise robustness reveal that model performance is governed by a balance between expressive capacity and trainability. These results provide empirical evidence that embedding quantum computational primitives within recurrent architectures offers a structurally grounded approach to improving sequence modeling capability. The proposed framework establishes a new design paradigm for integrating quantum operators into temporal learning models, with potential applications in complex dynamical system prediction tasks.

quant-ph

SSB-Based Sensing-Assisted Robust Beamforming for High-Mobility UAV Communications in LAWN

High-mobility uncrewed aerial vehicle (UAV) communications in low-altitude wireless networks (LAWN) demand reliable beamforming, while conventional feedback-based schemes suffer from excessive overhead and severe misalignment under rapid trajectory variations. To address this challenge, this paper proposes an SSB-based sensing-assisted predictive robust beamforming framework that replaces explicit channel state information (CSI) feedback with sensing-driven state estimation and uncertainty-aware optimization. Leveraging the periodic 'always-on' synchronization signal block (SSB), a hierarchical sensing algorithm tailored for hybrid digital-analog uniform planar arrays is developed, combining 2D range-velocity profiling and augmented beamspace multiple signal classification (MUSIC). By integrating a locally-focused analog receive beamformer, the proposed sensing design can ensure energy accumulates across different radio-frequency (RF) chains while resolving angular ambiguity. An extended Kalman filter (EKF) is further employed to track UAV states between sparse synchronization-signal (SS) bursts, and a covariance correction is introduced to characterize maneuver-induced prediction uncertainties. Based on the derived statistical distributions of range and angular parameters, the communication channel is modeled through predictive correlation matrices rather than instantaneous CSI, leading to a multi-user robust beamforming formulation that maximizes average network sum-rate under uncertainty. The resulting nonconvex problem is efficiently solved via successive convex approximation and alternating minimization. Simulation results demonstrate that the proposed framework significantly enhances spectral efficiency and link stability compared with feedback-based beamforming and non-robust beamforming design, particularly in high-mobility and large-SSB-interval scenarios.

eess.SP

Node-Based Soft-Output Fast Successive Cancellation List Decoding of Polar Codes

The soft-output successive cancellation list (SO-SCL) decoder provides a methodology for estimating the a-posteriori probability log-likelihood ratios by only leveraging the conventional SCL decoder of polar codes. However, the sequential decoding nature of SCL introduces high decoding latency to SO-SCL. In this paper, we incorporate node-based fast decoding into the SO-SCL framework. After addressing the challenge of soft output extraction in special node decoding, we proposed the soft-output fast SCL (SO-FSCL) decoding algorithm, along with its log-domain implementation and hardware-friendly version. The proposed SO-FSCL decoder can be regarded as an add-on extension to FSCL decoder, enabling us to autonomously choose whether to output only hard decisions like FSCL or to provide additional soft outputs. Latency and complexity analyses demonstrate that SO-FSCL can significantly reduce, for example, decoding time steps by 81.8\% (with unlimited resources), the number of additions by 41.3\%, and the number of comparisons by 46.4\%. Meanwhile, simulation results indicate that SO-FSCL delivers almost the same soft-output performance as SO-SCL, outperforming other soft-output polar decoders, especially in scenarios involving iterative decoding.

cs.IT

Low-Complexity Soft-Feedback Detector for AFDM Systems

Affine frequency division multiplexing (AFDM), an emerging multi-carrier modulation scheme, has garnered significant attention due to its resilience to Doppler shifts and capability to achieve full diversity in doubly dispersive channels. However, existing data detection algorithms for AFDM systems face a significant trade-off between computational complexity and accuracy. In this paper, a novel low-complexity data detection scheme, termed the soft-feedback detector (SFD), is proposed. Particularly, building upon a maximum ratio combining (MRC) estimator framework, the SFD leverages the a priori symbol distribution to mitigate error propagation during iterative detection. Specifically, soft-decision feedback is incorporated as extrinsic information derived from the log-likelihood ratios of the transmitted symbols. As a result, the proposed detector significantly enhances detection accuracy while maintaining low computational complexity. Simulation results demonstrate that the SFD consistently outperforms benchmark decision-feedback detectors. In particular, compared with the conventional MRC detector, the proposed scheme achieves approximately a 3 dB signal-to-noise ratio (SNR) gain at the bit error rate (BER) of $10^{-3}$.

eess.SP