SearcharxivSearch

arXiv subjects

Zhi Sun

Publications and source records attributed to Zhi Sun.

At least 19 recordsLinked to original sources

Physics Equivariance for Robust Generalization in Wireless Foundation Model

Wireless foundation models (WFMs) have recently emerged as a promising paradigm for learning multiple channel state information (CSI) acquisition tasks. However, unlike natural language tokens governed by statistical co-occurrence, wireless channels are generated by electromagnetic propagation laws, and current WFM training is constrained by limited data scale, narrow distribution coverage dominated by simulations, and a pronounced sim-to-real gap. As a result, simply scaling model parameters and CSI samples does not necessarily yield robust and generalizable models. In this paper, we advocate enabling physics equivariance as a principled and explainable inductive bias for WFMs. Specifically, we focus on a universal propagation property for electromagnetic waves, termed wave equivariance: when the input CSI is modulated along time-frequency-space dimensions, the output channel response should exhibit the corresponding transformation. Empirical studies show that the vanilla-WFM fails to reliably acquire such equivariance even with a large number of model parameters and training samples. To address this, we design the physics-intrinsic WFM (phys-WFM) with wave equivariance, which explicitly aligns model behaviors with an interpretable wave propagation structure. Results demonstrate that the proposed design effectively captures wave equivariance and substantially improves robustness and generalization to unseen environments under distribution shift, offering a physics-grounded and testable route toward explainable wireless foundation models.

eess.SP

Unified Generalization for Frequency-Domain Channel Extrapolation Across Near-Field and Far-Field Scenarios

As antenna arrays grow, near-field effects become non-negligible in large-scale MIMO, making accurate low-overhead channel acquisition crucial in both far-field and near-field regimes. Deep-learning-based frequency-domain channel extrapolation can reduce pilot overhead, but existing extrapolators generalize poorly to unseen distances and environments, especially across near-field and far-field channels. We propose a physically interpretable framework to unify generalization across both regimes. Our key insight is that angular profiles are regime-dependent, while delay profiles share a sparsity structure that can be aligned. Based on this, we develop a physics-guided disentanglement and alignment pipeline with multi-cluster decoupling, angle-delay feature disentanglement, and delay-domain alignment, enabling the model to learn distribution-stable delay features while reusing heterogeneous angular features. We further design a unified near/far-field DL extrapolator (UNiFi-DLE) and detail its dataset preparation, training, and inference. Simulations and sim-to-real experiments show that UNiFi-DLE generalizes robustly to unseen near-field and far-field scenarios and consistently outperforms state-of-the-art methods.

eess.SP

Hyperbolic Frequency Multicarrier Modulation for Wideband Linear Time-Varying Channels

Numerous multicarrier modulation schemes have been proposed recently to enhance the performance in narrowband doubly dispersive channels for emerging high-mobility applications. However, the ultra-reliable modulation framework in wideband linear time-varying (LTV) channels remains an open problem, where the time dilations and contractions brought by the high mobility cannot be ignored for the baseband signal to obtain the constant Doppler shift across the whole transmission band. To solve this problem, we propose the hyperbolic frequency multicarrier (HFMC) waveform in this paper based on the inspiration from affine frequency division multiplexing (AFDM) modulation, where the delay and Doppler shift are absorbed into a 1D shift in the affine domain to provide a compact characterization of doubly dispersive discrete-time channels. By adopting the passband representation of wideband LTV channels and hyperbolic frequency modulated (HFM) signals, we reveal that the Doppler scaling factor brought by the relative mobility can be absorbed into an equivalent delay. The basic principle of HFMC modulation is established by investigating the approximate orthogonality among HFMC subcarriers, which are generated from a basic HFM signal by utilizing uniformly spaced equivalent delay. The spectrum of HFMC subcarriers is also analyzed to evaluate the system capacity, where the overlapping nature in the frequency domain can be observed. The input-output characterization in wideband LTV channels is then executed to confirm the 1D integration of time delay and Doppler scaling factor for each path, which demonstrates the ability to exploit potential multipath diversity. The parameter optimization based on the input-output relation and spectrum analysis is finally developed to balance the efficiency and reliability.

eess.SP

KARMA: Knowledge-Action Regularized Multimodal Alignment for Personalized Search at Taobao

Large Language Models (LLMs) are equipped with profound semantic knowledge, making them a natural choice for injecting semantic generalization into personalized search systems. However, in practice we find that directly fine-tuning LLMs on industrial personalized tasks (e.g. next item prediction) often yields suboptimal results. We attribute this bottleneck to a critical Knowledge--Action Gap: the inherent conflict between preserving pre-trained semantic knowledge and aligning with specific personalized actions by discriminative objectives. Empirically, action-only training objectives induce Semantic Collapse, such as attention "sinks". This degradation severely cripples the LLM's generalization, failing to bring improvements to personalized search systems. We propose KARMA (Knowledge--Action Regularized Multimodal Alignment), a unified framework that treats semantic reconstruction as a train-only regularizer. KARMA optimizes a next-interest embedding for retrieval (Action) while enforcing semantic decodability (Knowledge) through two complementary objectives: (i) history-conditioned semantic generation, which anchors optimization to the LLM's native next-token distribution, and (ii) embedding-conditioned semantic reconstruction, which constrains the interest embedding to remain semantically recoverable. On Taobao search system, KARMA mitigates semantic collapse (attention-sink analysis) and improves both action metrics and semantic fidelity. In ablations, semantic decodability yields up to +22.5 HR@200. With KARMA, we achieve +0.25 CTR AUC in ranking, +1.86 HR in pre-ranking and +2.51 HR in recalling. Deployed online with low inference overhead at ranking & pre-ranking stage, KARMA drives +0.9% increase in GMV.

cs.IR

Generalizable Learning for Massive MIMO CSI Feedback in Unseen Environments

Deep learning is promising to enhance the accuracy and reduce the overhead of channel state information (CSI) feedback, which can boost the capacity of frequency division duplex (FDD) massive multiple-input multiple-output (MIMO) systems. Nevertheless, the generalizability of current deep learning-based CSI feedback algorithms cannot be guaranteed in unseen environments, which induces a high deployment cost. In this paper, the generalizability of deep learning-based CSI feedback is promoted with physics interpretation. Firstly, the distribution shift of the cluster-based channel is modeled, which comprises the multi-cluster structure and single-cluster response. Secondly, the physics-based distribution alignment is proposed to effectively address the distribution shift of the cluster-based channel, which comprises multi-cluster decoupling and fine-grained alignment. Thirdly, the efficiency and robustness of physics-based distribution alignment are enhanced. Explicitly, an efficient multi-cluster decoupling algorithm is proposed based on the Eckart-Young-Mirsky (EYM) theorem to support real-time CSI feedback. Meanwhile, a hybrid criterion to estimate the number of decoupled clusters is designed, which enhances the robustness against channel estimation error. Fourthly, environment-generalizable neural network for CSI feedback (EG-CsiNet) is proposed as a novel learning framework with physics-based distribution alignment. Based on extensive simulations and sim-to-real experiments in various conditions, the proposed EG-CsiNet can robustly reduce the generalization error by more than 3 dB compared to the state-of-the-arts.

eess.SP

Acoustic RIS for Massive Spatial Multiplexing: Unleashing Degrees of Freedom and Capacity in Underwater Communications

Underwater acoustic (UWA) communications are essential for high-speed marine data transmission but remain severely constrained by limited bandwidth, significant propagation loss, and sparse multipath structures. Conventional underwater acoustic multiple-input multiple-output (MIMO) systems primarily utilize spatial diversity but suffer from limited array resolution, causing angular ambiguity and insufficient spatial degrees of freedom (DoFs). This paper addresses these limitations through acoustic Reconfigurable Intelligent Surfaces (aRIS) to actively generate orthogonally distinguishable virtual paths, significantly enhancing spatial DoFs and channel capacity. An ocean-specific DoF-channel coupling model is established, explicitly deriving conditions for spatial rank enhancement. Subsequently, the optimal geometric locus, termed the Light-Point, is analytically identified, where deploying a single aRIS maximizes DoFs by introducing two and three additional resolvable paths in deep-sea and shallow-sea environments, respectively. Furthermore, an active simultaneous transmitting and reflecting (ASTAR) aRIS architecture with independent beam control and adaptive beam-tracking mechanism integrating unmanned underwater vehicles (UUVs) and acoustic intensity gradient sensing is proposed. Extensive simulations validate the proposed joint aRIS deployment and beamforming framework, demonstrating substantial UWA channel capacity improvements-up to 265% and 170% in shallow-sea and deep-sea scenarios, respectively.

cs.NI

On the Characterization and Evaluation of Doppler Squint in Wideband ODDM Systems

The recently proposed orthogonal delay-Doppler division multiplexing (ODDM) modulation has been demonstrated to enjoy excellent reliability over doubly-dispersive channels. However, most of the prior analysis tends to ignore the interactive dispersion caused by the wideband property of ODDM signal, which possibly leads to performance degradation. To solve this problem, we investigate the input-output relation of ODDM systems considering the wideband effect, which is also known as the Doppler squint effect (DSE) in the literature. The extra delay-Doppler (DD) dispersion caused by the DSE is first explicitly explained by employing the time-variant frequency response of multipath channels. Its characterization is then derived for both reduced cyclic prefix (RCP) and zero padded (ZP)-based wideband ODDM systems, where the extra DD spread and more complicated power leakage outside the peak region are presented theoretically. Numerical results are finally provided to confirm the significance of DSE. The derivations in this paper are beneficial for developing accurate signal processing techniques in ODDM-based integrated sensing and communication systems.

eess.SP

Enhancing Environment Generalizability for Deep Learning-Based CSI Feedback

Accurate and low-overhead channel state information (CSI) feedback is essential to boost the capacity of frequency division duplex (FDD) massive multiple-input multiple-output (MIMO) systems. Deep learning-based CSI feedback significantly outperforms conventional approaches. Nevertheless, current deep learning-based CSI feedback algorithms exhibit limited generalizability to unseen environments, which obviously increases the deployment cost. In this paper, we first model the distribution shift of CSI across different environments, which is composed of the distribution shift of multipath structure and a single-path. Then, EG-CsiNet is proposed as a novel CSI feedback learning framework to enhance environment-generalizability. Explicitly, EG-CsiNet comprises the modules of multipath decoupling and fine-grained alignment, which can address the distribution shift of multipath structure and a single path. Based on extensive simulations, the proposed EG-CsiNet can robustly enhance the generalizability in unseen environments compared to the state-of-the-art, especially in challenging conditions with a single source environment.

eess.SP

Data-Driven Modulation Optimization with LMMSE Equalization for Reliability Enhancement in Underwater Acoustic Communications

Ultra-reliable underwater acoustic (UWA) communications serve as one of the key enabling technologies for future space-air-ground-underwater integrated networks. However, the reliability of current UWA transmission is still insufficient since severe performance degradation occurs for conventional multicarrier systems in UWA channels with severe delay-scale spread. To solve this problem, we exploit learning-inspired approaches to optimize the modulation scheme under the assumption of linear minimum mean square error (LMMSE) equalization, where the discrete representation of waveforms is adopted by utilizing Nyquist filters. The optimization problem is first transferred into maximizing the fairness of estimation mean square error (MSE) for each data symbol since the total MSE is invariant considering the property of orthogonal modulation. The Siamese architecture is then adopted to obtain consistent optimization results across various channel conditions, which avoids the overhead of online feedback, cooperation, and deployment of neural networks and guarantees generalization. The overall scheme including the loss function, neural network structure, and training process is also investigated in depth in this paper. The excellent performance and robustness of the proposed modulation scheme are verified by carrying out the bit error rate test over various UWA channels with severe delay-scale spread.

eess.SP

Generalizable Learning for Frequency-Domain Channel Extrapolation under Distribution Shift

Frequency-domain channel extrapolation is effective in reducing pilot overhead for massive multiple-input multiple-output (MIMO) systems. Recently, Deep learning (DL) based channel extrapolator has become a promising candidate for modeling complex frequency-domain dependency. Nevertheless, current DL extrapolators fail to operate in unseen environments under distribution shift, which poses challenges for large-scale deployment. In this paper, environment generalizable learning for channel extrapolation is achieved by realizing distribution alignment from a physics perspective. Firstly, the distribution shift of wireless channels is rigorously analyzed, which comprises the distribution shift of multipath structure and single-path response. Secondly, a physics-based progressive distribution alignment strategy is proposed to address the distribution shift, which includes successive path-oriented design and path alignment. Path-oriented DL extrapolator decomposes multipath channel extrapolation into parallel extrapolations of the extracted path, which can mitigate the distribution shift of multipath structure. Path alignment is proposed to address the distribution shift of single-path response in path-oriented DL extrapolators, which eventually enables generalizable learning for channel extrapolation. In the simulation, distinct wireless environments are generated using the precise ray-tracing tool. Based on extensive evaluations, the proposed path-oriented DL extrapolator with path alignment can reduce extrapolation error by more than 6 dB in unseen environments compared to the state-of-the-arts.

eess.SP

Path Evolution Model for Endogenous Channel Digital Twin towards 6G Wireless Networks

Massive Multiple Input Multiple Output (MIMO) is critical for boosting 6G wireless network capacity. Nevertheless, high dimensional Channel State Information (CSI) acquisition becomes the bottleneck of 6G massive MIMO system. Recently, Channel Digital Twin (CDT), which replicates physical entities in wireless channels, has been proposed, providing site-specific prior knowledge for CSI acquisition. However, external devices (e.g., cameras and GPS devices) cannot always be integrated into existing communication systems, nor are they universally available across all scenarios. Moreover, the trained CDT model cannot be directly applied in new environments, which lacks environmental generalizability. To this end, Path Evolution Model (PEM) is proposed as an alternative CDT to reflect physical path evolutions from consecutive channel measurements. Compared to existing CDTs, PEM demonstrates virtues of full endogeneity, self-sustainability and environmental generalizability. Firstly, PEM only requires existing channel measurements, which is free of other hardware devices and can be readily deployed. Secondly, self-sustaining maintenance of PEM can be achieved in dynamic channel by progressive updates. Thirdly, environmental generalizability can greatly reduce deployment costs in dynamic environments. To facilitate the implementation of PEM, an intelligent and light-weighted operation framework is firstly designed. Then, the environmental generalizability of PEM is rigorously analyzed. Next, efficient learning approaches are proposed to reduce the amount of training data practically. Extensive simulation results reveal that PEM can simultaneously achieve high-precision and low-overhead CSI acquisition, which can serve as a fundamental CDT for 6G wireless networks.

eess.SP

Covering Underwater Shadow Zones using Acoustic Reconfigurable Intelligent Surfaces

To better explore the oceans, seamless communication coverage of the vast 3D underwater space is desired. Unlike terrestrial networks using radio signals, underwater acoustic communications face a unique challenge: nodes in underwater shadow zones cannot connect to the network, even within the line of sight. These shadow zones can extend for tens of kilometers, causing communication nodes to disconnect. Existing efforts focus on passive avoidance of shadow zones, but this strategy cannot ensure seamless coverage in dynamic ocean environments. This paper addresses the shadow zone problem by utilizing acoustic Reconfigurable Intelligent Surfaces (RIS) to actively control the underwater channel. Shadow zones are analytically modeled, and optimal RIS deployment strategies are developed for both deep-sea and shallow-sea environments. The acoustic RIS is redesigned considering practical engineering limitations and validated through pool tests. Bellhop-based simulations show that without RIS deployment, coverage is limited to less than 20%, regardless of source strength. However, with optimal RIS deployment, energy coverage can reach almost 100%.

cs.NI

On-demand Quick Metasurface Design with Neighborhood Attention Transformer

Metasurfaces are reshaping traditional optical paradigms and are increasingly required in complex applications that demand substantial computational resources to numerically solve Maxwell's equations-particularly for large-scale systems, inhomogeneous media, and densely packed metadevices. Conventional forward design using electromagnetic solvers is based on specific approximations, which may not effectively address complex problems. In contrast, existing inverse design methods are a stepwise process that is often time-consuming and involves repetitive computations. Here, we present an inverse design approach utilizing a surrogate Neighborhood Attention Transformer, MetaE-former, to predict the performance of metasurfaces with ultrafast speed and high accuracy. This method achieves global solutions for hundreds of nanostructures simultaneously, providing up to a 250,000-fold speedup compared with solving for individual meta-atoms based on the FDTD method. As examples, we demonstrate a binarized high-numerical-aperture (about 1.31) metalens and several optimized structured-light meta-generators. Our method significantly improves the beam shaping adaptability with metasurfaces and paves the way for fast designing of large-scale metadevices for shaping extreme light fields with high accuracy.

physics.optics

Analysis and Optimization of Multiple-STAR-RIS Assisted MIMO-NOMA with GSVD Precoding: An Operator-Valued Free Probability Approach

Among the key enabling 6G techniques, multiple-input multiple-output (MIMO) and non-orthogonal multiple-access (NOMA) play an important role in enhancing the spectral efficiency of the wireless communication systems. To further extend the coverage and the capacity, the simultaneously transmitting and reflecting reconfigurable intelligent surface (STAR-RIS) has recently emerged out as a cost-effective technology. To exploit the benefit of STAR-RIS in the MIMO-NOMA systems, in this paper, we investigate the analysis and optimization of the downlink dual-user MIMO-NOMA systems assisted by multiple STAR-RISs under the generalized singular value decomposition (GSVD) precoding scheme, in which the channel is assumed to be Rician faded with the Weichselberger's correlation structure. To analyze the asymptotic information rate of the users, we apply the operator-valued free probability theory to obtain the Cauchy transform of the generalized singular values (GSVs) of the MIMO-NOMA channel matrices, which can be used to obtain the information rate by Riemann integral. Then, considering the special case when the channels between the BS and the STAR-RISs are deterministic, we obtain the closed-form expression for the asymptotic information rates of the users. Furthermore, a projected gradient ascent method (PGAM) is proposed with the derived closed-form expression to design the STAR-RISs thereby maximizing the sum rate based on the statistical channel state information. The numerical results show the accuracy of the asymptotic expression compared to the Monte Carlo simulations and the superiority of the proposed PGAM algorithm.

eess.SP

Augmenting Channel Simulator and Semi- Supervised Learning for Efficient Indoor Positioning

This work aims to tackle the labor-intensive and resource-consuming task of indoor positioning by proposing an efficient approach. The proposed approach involves the introduction of a semi-supervised learning (SSL) with a biased teacher (SSLB) algorithm, which effectively utilizes both labeled and unlabeled channel data. To reduce measurement expenses, unlabeled data is generated using an updated channel simulator (UCHS), and then weighted by adaptive confidence values to simplify the tuning of hyperparameters. Simulation results demonstrate that the proposed strategy achieves superior performance while minimizing measurement overhead and training expense compared to existing benchmarks, offering a valuable and practical solution for indoor positioning.

eess.SP

GaitSADA: Self-Aligned Domain Adaptation for mmWave Gait Recognition

mmWave radar-based gait recognition is a novel user identification method that captures human gait biometrics from mmWave radar return signals. This technology offers privacy protection and is resilient to weather and lighting conditions. However, its generalization performance is yet unknown and limits its practical deployment. To address this problem, in this paper, a non-synthetic dataset is collected and analyzed to reveal the presence of spatial and temporal domain shifts in mmWave gait biometric data, which significantly impacts identification accuracy. To mitigate this issue, a novel self-aligned domain adaptation method called GaitSADA is proposed. GaitSADA improves system generalization performance by using a two-stage semi-supervised model training approach. The first stage employs semi-supervised contrastive learning to learn a compact gait representation from both source and target domain data, aligning source-target domain distributions implicitly. The second stage uses semi-supervised consistency training with centroid alignment to further close source-target domain gap by pseudo-labelling the target-domain samples, clustering together the samples belonging to the same class but from different domains, and pushing the class centroid close to the weight vector of each class. Experiments show that GaitSADA outperforms representative domain adaptation methods with an improvement ranging from 15.41\% to 26.32\% on average accuracy in low data regimes. Code and dataset will be available at https://exitudio.github.io/GaitSADA

cs.CV

On the Mutual Information of Multi-RIS Assisted MIMO: From Operator-Valued Free Probability Aspect

The reconfigurable intelligent surface (RIS) is useful to effectively improve the coverage and data rate of end-to-end communications. In contrast to the well-studied coverage-extension use case, in this paper, multiple RIS panels are introduced, aiming to enhance the data rate of multi-input multi-output (MIMO) channels in presence of insufficient scattering. Specifically, via the operator-valued free probability theory, the asymptotic mutual information of the large-dimensional RIS-assisted MIMO channel is obtained under the Rician fading with Weichselberger's correlation structure, in presence of both the direct and the reflected links. Although the mutual information of Rician MIMO channels scales linearly as the number of antennas and the signal-to-noise ratio (SNR) in decibels, numerical results show that it requires sufficiently large SNR, proportional to the Rician factor, in order to obtain the theoretically guaranteed linear improvement. This paper shows that the proposed multi-RIS deployment is especially effective to improve the mutual information of MIMO channels under the large Rician factor conditions. When the reflected links have similar arriving and departing angles across the RIS panels, a small number of RIS panels are sufficient to harness the spatial degree of freedom of the multi-RIS assisted MIMO channels.

cs.IT

Chirp-based Hierarchical Beam Training for Extremely Large-Scale Massive MIMO

XL-MIMO promises to provide ultrahigh data rates in Terahertz (THz) spectrum. However, the spherical-wavefront wireless transmission caused by large aperture array presents huge challenges for channel state information (CSI) acquisition. Two independent parameters (physical angles and transmission distance) should be simultaneously considered in XL-MIMO beamforming, which brings severe overhead consumption and beamforming degradation. To address this problem, we exploit the near-field channel characteristic and propose one low-overhead hierarchical beam training scheme for near-field XL-MIMO system. Firstly, we project near-field channel into spatial-angular domain and slope-intercept domain to capture detailed representations. Secondly, a novel spatial-chirp beam-aided codebook and corresponding hierarchical update policy are proposed. Theoretical analyses and numerical simulations are also displayed to verify the superior performances on beamforming and training overhead.

cs.IT