Searcharxiv⌕ Search

arXiv subjects

Xiaojuan Zhang

Publications and source records attributed to Xiaojuan Zhang.

11 recordsLinked to original sources

Visibility-Region Coupling in XL-MIMO AGV Fleets: Triple-Role Modeling and Masked Beamforming

Extremely large-scale multiple-input multiple-output (XL-MIMO) is a promising technology for supporting automated guided vehicle (AGV) fleets in smart port terminals. However, the metallic container environment induces spatial non-stationarity, whereby each AGV is visible to only a subset of the array, referred to as its visibility region (VR). Unlike existing XL-MIMO models that assume user-independent VRs, we show that each AGV simultaneously acts as a communication user, a metallic scatterer, and a blocker, resulting in coupled user channels and VRs. We formulate this \emph{triple-role} effect through a VR-coupled channel model and develop a VR-aware downlink beamforming framework based on masked weighted minimum mean-square error (WMMSE), where the masking operation exactly enforces VR support constraints while significantly reducing computational complexity. Simulation results in a realistic smart port scenario demonstrate more than a threefold sum-rate improvement over VR-unaware baselines, with the gains becoming increasingly pronounced as fleet density increases.

cs.IT↗

A Unified Dual Framework for Sparse-Array Near-Field Beam Focusing With Spatial Interference Suppression

We study sparse-array near-field beam focusing with spatial interference suppression, a problem arising in coherent satellite formations and other distributed non-terrestrial arrays. State-of-the-art designs solve it numerically through second-order cone programming (SOCP) with cutting-plane refinement, yet the achievable signal-to-interference ratio (SIR) and its link to classical adaptive beamforming have remained without an analytical characterization. We supply this characterization via a Lagrangian-dual analysis, obtaining three results. First, every optimal beamformer is a generalized matched filter against an effective spatial covariance induced by an optimal dual measure; this closed form recovers MVDR, LCMV, and SOCP-based focusing as special cases. Second, the dual measure has finite support of cardinality at most $M^2$ in general, sharpening to $M$ for uniform linear arrays ($M$ the number of array elements), which yields a finite-dimensional convergence certificate for cutting-plane methods. Third, a closed-form upper bound on the mean-SIR admits an asymptotic logarithmic scaling law in $M$ under near-collinear geometry, identifying array order, rather than the optimization algorithm, as the dominant performance factor. A Riemannian conjugate-gradient algorithm on the unit-torus manifold is developed for practical constant-modulus beamforming, and numerical results demonstrate that it closely approaches the derived performance limit.

cs.IT↗

Twin-in-the-Loop Optimization and Fundamental Limits of Position--Velocity Estimation in Cell-Free ISAC Systems

Digital twin (DT) networks require tight integration with wireless sensing, yet the fundamental limits of such coupling in cell-free integrated sensing and communication (ISAC) systems remain largely unexplored, particularly in the presence of fluid intelligent metasurfaces (FIM). This paper establishes a joint position-velocity Cramer-Rao bound (CRB) framework, operationalized through a twin-in-the-loop architecture. By leveraging a scatter-matrix decomposition of the velocity Fisher information, we show that single-base-station systems are inherently rank-deficient for two-dimensional velocity estimation, whereas cell-free deployments with multiple access-point pairs achieve full observability. The resulting CRB reveals a spatio-temporal decoupling: FIM shape optimization significantly improves position accuracy but does not affect the velocity CRB under isotropic waveforms while Doppler coupling asymmetrically enhances position estimation accuracy. Building on this analysis, we develop a closed-loop DT framework, deriving the critical mismatch angle in closed form and showing that angular diversity in cell-free systems mitigates DT prediction errors. We further characterize the optimal synchronization period and propose a confidence-aware scheduling strategy that reduces the DT update rate. Numerical results demonstrate substantial performance gains over single-base-station systems, with improvements attributed to angular diversity, Doppler-position coupling, and FIM adaptation.

cs.IT↗

Risk-Aware Safe Throughput Forecasting for Starlink Networks

As a representative low Earth orbit (LEO) broadband system, Starlink exhibits highly variable access throughput, making short-term forecasting essential for network resource management. Existing forecasting methods mainly optimize symmetric point-prediction metrics such as MAE and RMSE, but they do not explicitly control the asymmetric risk of overestimating future throughput, which can cause over-admission, bandwidth overbooking, and service violations. This paper formulates Starlink throughput prediction as a risk-budgeted safe forecasting problem, where the predictor must satisfy a prescribed overestimation budget while maintaining competitive accuracy. We propose Budget-Guided Coarse-to-Fine Quantile Selection (BG-CFQS), a data-driven framework that trains a family of lower-quantile predictors, locates the quantile boundary satisfying the risk budget, and refines the boundary region to select the most accurate feasible predictor. Experiments on three real-world Starlink throughput datasets show that BG-CFQS satisfies the risk budget on all datasets and achieves the lowest average MAE, mean positive error, and tail positive error among budget-feasible methods. In high-risk and severe-risk low-throughput regimes, BG-CFQS reduces harmful positive errors by 11.0% and 12.6%, respectively. An admission-control evaluation further shows that the proposed safe forecasts reduce dropped sessions, demonstrating that risk-aware forecasting can translate prediction safety into application-level benefits.

eess.SY↗

DiffPixelFormer: Differential Pixel-Aware Transformer for RGB-D Indoor Scene Segmentation

Indoor semantic segmentation is fundamental to computer vision and robotics, supporting applications such as autonomous navigation, augmented reality, and smart environments. Although RGB-D fusion leverages complementary appearance and geometric cues, existing methods often depend on computationally intensive cross-attention mechanisms and insufficiently model intra- and inter-modal feature relationships, resulting in imprecise feature alignment and limited discriminative representation. To address these challenges, we propose DiffPixelFormer, a differential pixel-aware Transformer for RGB-D indoor scene segmentation that simultaneously enhances intra-modal representations and models inter-modal interactions. At its core, the Intra-Inter Modal Interaction Block (IIMIB) captures intra-modal long-range dependencies via self-attention and models inter-modal interactions with the Differential-Shared Inter-Modal (DSIM) module to disentangle modality-specific and shared cues, enabling fine-grained, pixel-level cross-modal alignment. Furthermore, a dynamic fusion strategy balances modality contributions and fully exploits RGB-D information according to scene characteristics. Extensive experiments on the SUN RGB-D and NYUDv2 benchmarks demonstrate that DiffPixelFormer-L achieves mIoU scores of 54.28% and 59.95%, outperforming DFormer-L by 1.78% and 2.75%, respectively. Code is available at https://github.com/gongyan1/DiffPixelFormer.

cs.CV↗

OTFS for Joint Radar and Communication: Algorithms, Prototypes, and Experiments

We propose an Joint Radar and Communication (JRC) system that utilizes the Orthogonal Time Frequency Space (OTFS) signals. The system features a fast radar sensing algorithm for detecting target range and speed by using the OTFS communication signals, and a self-interference cancellation for enhanced multi-target separation. In addition to target detection, we propose methods for monitoring human vital signs, such as breathing rate and heartbeat. Furthermore, we explore two approaches for distinguishing between human and nonhuman targets: one based on signal processing and the other based on machine learning. We have developed a prototype JRC system using the software-defined radio (SDR) technology. Experimental results are shown to demonstrate the effectiveness of the prototype in detecting range, speed, and vital signs in both human and mobile robot scenarios, as well as in distinguishing between human and non-human targets.

cs.IT↗

Coordinated FMCW and OFDM for Integrated Sensing and Communication

We propose a coordinated FMCW-OFDM (Co-FMCW-OFDM) system that enables integrated sensing and communication (ISAC) by allowing sensing and communication to share the same RF front end, antennas, and spectral resources. In the proposed ISAC system, the FMCW signal is superimposed on the OFDM signal and serves dual purposes: facilitating bistatic sensing and enabling channel estimation at the receiver end. Based on proposed Co-FMCW-OFDM waveform, we propose two efficient sensing algorithms-fast cyclic correlation radar (FCCR) and digital mixing and down-sampling (DMD)- which significantly reduce system complexity while accurately estimating target range and velocity. We consider a realistic channel model where delays can take any value, not just integer multiples of the sampling period. This leads to a significantly larger number of effective paths compared to the actual number of targets, which makes the sensing, channel estimation, and interference cancellation more challenging. Leveraging the sensing results, we develop a sensing-aided effective channel estimation method which effectively reconstructs the channel under arbitrary delay condition based on successive interference cancellation and propose an interference cancellation scheme that removes the FMCW signal before the OFDM demodulation. Simulation results demonstrate that the proposed system achieves superior sensing accuracy, improved channel estimation, and lower bit error rate (BER) compared to conventional OFDM systems with embedded pilots. The proposed scheme demonstrates superior BER performance in comparison to the conventional OFDM-plus-FMCW approach.

cs.IT↗

Progressive Bird's Eye View Perception for Safety-Critical Autonomous Driving: A Comprehensive Survey

Bird's-Eye-View (BEV) perception has become a foundational paradigm in autonomous driving, enabling unified spatial representations that support robust multi-sensor fusion and multi-agent collaboration. As autonomous vehicles transition from controlled environments to real-world deployment, ensuring the safety and reliability of BEV perception in complex scenarios - such as occlusions, adverse weather, and dynamic traffic - remains a critical challenge. This survey provides the first comprehensive review of BEV perception from a safety-critical perspective, systematically analyzing state-of-the-art frameworks and implementation strategies across three progressive stages: single-modality vehicle-side, multimodal vehicle-side, and multi-agent collaborative perception. Furthermore, we examine public datasets encompassing vehicle-side, roadside, and collaborative settings, evaluating their relevance to safety and robustness. We also identify key open-world challenges - including open-set recognition, large-scale unlabeled data, sensor degradation, and inter-agent communication latency - and outline future research directions, such as integration with end-to-end autonomous driving systems, embodied intelligence, and large language models.

cs.RO↗

AI-Based Impedance Encoding-Decoding Method for Online Impedance Network Construction of Wind Farms

The impedance network (IN) model is gaining popularity in the oscillation analysis of wind farms. However, the construction of such an IN model requires impedance curves of each wind turbine under their respective operating conditions, making its online application difficult due to the transmission of numerous high-density impedance curves. To address this issue, this paper proposes an AI-based impedance encoding-decoding method to facilitate the online construction of IN model. First, an impedance encoder is trained to compress impedance curves by setting the number of neurons much smaller than that of frequency points. Then, the compressed data of each turbine are uploaded to the wind farm and an impedance decoder is trained to reconstruct original impedance curves. At last, based on the nodal admittance matrix (NAM) method, the IN model of the wind farm can be obtained. The proposed method is validated via model training and real-time simulations, demonstrating that the encoded impedance vectors enable fast transmission and accurate reconstruction of the original impedance curves.

eess.SP↗

Variational Disentangled Graph Auto-Encoders for Link Prediction

With the explosion of graph-structured data, link prediction has emerged as an increasingly important task. Embedding methods for link prediction utilize neural networks to generate node embeddings, which are subsequently employed to predict links between nodes. However, the existing embedding methods typically take a holistic strategy to learn node embeddings and ignore the entanglement of latent factors. As a result, entangled embeddings fail to effectively capture the underlying information and are vulnerable to irrelevant information, leading to unconvincing and uninterpretable link prediction results. To address these challenges, this paper proposes a novel framework with two variants, the disentangled graph auto-encoder (DGAE) and the variational disentangled graph auto-encoder (VDGAE). Our work provides a pioneering effort to apply the disentanglement strategy to link prediction. The proposed framework infers the latent factors that cause edges in the graph and disentangles the representation into multiple channels corresponding to unique latent factors, which contributes to improving the performance of link prediction. To further encourage the embeddings to capture mutually exclusive latent factors, we introduce mutual information regularization to enhance the independence among different channels. Extensive experiments on various real-world benchmarks demonstrate that our proposed methods achieve state-of-the-art results compared to a variety of strong baselines on link prediction tasks. Qualitative analysis on the synthetic dataset also illustrates that the proposed methods can capture distinct latent factors that cause links, providing empirical evidence that our models are able to explain the results of link prediction to some extent. All code will be made publicly available upon publication of the paper.

cs.LG↗

Contrastive Disentangled Learning on Graph for Node Classification

Contrastive learning methods have attracted considerable attention due to their remarkable success in analyzing graph-structured data. Inspired by the success of contrastive learning, we propose a novel framework for contrastive disentangled learning on graphs, employing a disentangled graph encoder and two carefully crafted self-supervision signals. Specifically, we introduce a disentangled graph encoder to enforce the framework to distinguish various latent factors corresponding to underlying semantic information and learn the disentangled node embeddings. Moreover, to overcome the heavy reliance on labels, we design two self-supervision signals, namely node specificity and channel independence, which capture informative knowledge without the need for labeled data, thereby guiding the automatic disentanglement of nodes. Finally, we perform node classification tasks on three citation networks by using the disentangled node embeddings, and the relevant analysis is provided. Experimental results validate the effectiveness of the proposed framework compared with various baselines.

cs.LG↗