Searcharxiv⌕ Search

arXiv subjects

Yingzhuang Liu

Publications and source records attributed to Yingzhuang Liu.

15 recordsLinked to original sources

Beyond the Matrix Sign: Quadratic Spectral Descent

Muon emerges as a strong competitor of the AdamW for LLM pretraining, because the matrix-wise update it employs can potentially incur smaller second-order penalty than the once dominating AdamW, which performs coordinate-wise update. However, the spectral flattening procedure in Muon is quite debatable since it discards the spectral amplitude information totally. To seek for better spectral allocation (and the associated spectral subspace), we propose to solve the quadratic model of loss function under the spectral norm constraint \textit{directly} (i.e., in a genuinely Newtonian way) and thus obtaining the Quadratic Spectral Descent (QSD) algorithm. In contrast, many existing curvature-aware methods either exploit the second-order information in an \textit{implicit} way by changing the weight update geometry (such as Mousse, FISMO) or rely on strong assumptions (such as the weight displacement isotropy assumption in Newton-Muon). QSD's potential advantage over these methods is best illustrated in the isotropic curvature scenario, where Mousse, FISMO and Newton-Muon all reduce to Muon while the spectral allocation in QSD is still \textit{non-flat} (since the spectral allocation in QSD depends on the \textit{gradient to curvature ratio}). Meanwhile, to control the complexity of QSD, we employ inversion-free K-FAC and \textit{online} Frank-Wolfe update which is essentially a matrix sign operator. Overall, the complexity increase can be rather mild. Experiments on GPT pre-training show that QSD consistently improves validation loss over Muon and recent Muon variants, while achieving up to an $8.49\%$ wall-clock speedup at matched validation loss.

cs.LG↗

Digital Self-Interference Cancellation in Full-Duplex Radios: A Fundamental Limit Perspective

D-SIC is of crucial importance for the implementation of IBFD radios. Unfortunately, the achievable performance limit remains underexplored. To fill this gap, in this paper we aim to explore the performance limit, i.e., the minimum residual self-interference (RSI) of the most commonly used PH canceller, and provide the achievable pilot design accordingly. To this end, we first conduct a systematic analysis of the RSI power for the PH canceller, which takes into account both the truncation-induced error and the noise-induced error, whereas the former is usually ignored in the existing works. To simplify the performance analysis of RSI power, we employ the generalized Laguerre polynomial (GLP)-based PH canceller instead of the conventional monomial-based one, due to the appealing orthogonality property of the GLP for Gaussian inputs. With the GLP representation of the PH canceller, we further prove that the least-squares channel estimator is asymptotically unbiased, thus demonstrating the asymptotic optimality of Gaussian pilot sequences. Moreover, for the pilot sequence with a finite length, a succinct criterion for minimizing the RSI, namely, the condition-number-to-minimum eigenvalue ratio (CMER) criterion, which essentially balances the truncation-induced and noise-induced error, is presented. By contrast, the existing works normally consider the latter only. Interestingly, it is revealed that an appropriate PAPR of the pilot sequence is of critical importance to achieve the above balance. Simulation results demonstrate that the pilot sequence optimized according to our proposed CMER criterion can achieve an RSI as low as -87.3 dBm, which is over 14 dB lower than that of HE-LTF and over 6 dB lower than that of the state-of-the-art pilot sequence proposed in [1], provided that the order of the PH canceller is no higher than 9 because of the complexity constraint.

eess.SP↗

Rényi Sharpness: A Novel Sharpness that Strongly Correlates with Generalization

Sharpness (of the loss minima) is widely believed to be a good indicator of generalization of neural networks. Unfortunately, the correlation between existing sharpness measures and generalization is not as strong as expected, and sometimes even contradiction occurs. To address this problem, a key observation in this paper is: what really matters for generalization is the average spread (or unevenness) of the spectrum of loss Hessian $\mathbf{H}$. For this reason, conventional sharpness measures, such as trace sharpness $\operatorname{tr}(\mathbf{H})$, which cares about the average value of the spectrum, or max-eigenvalue sharpness $λ_{\max}(\mathbf{H})$, which concerns the maximum spread of the spectrum, are not sufficient to well predict generalization. To characterize the average spread of the Hessian spectrum, we leverage the notion of Rényi entropy in information theory, which captures the unevenness of a probability vector and can thus be extended to a general non-negative vector, such as the Hessian spectrum at loss minima. Specifically, we propose Rényi sharpness, defined as the negative of the Rényi entropy of loss Hessian $\mathbf{H}$. Extensive experiments demonstrate that Rényi sharpness exhibits strong and consistent correlation with generalization in various scenarios. Moreover, two generalization bounds with respect to Rényi sharpness are established by exploiting its desirable reparametrization invariance property. Finally, as an initial attempt to exploit Rényi sharpness for regularization, Rényi Sharpness Aware Minimization (RSAM) is proposed, where a variant of Rényi sharpness is used as the regularizer. RSAM is competitive with state-of-the-art SAM algorithms and far better than conventional SAM based on max-eigenvalue sharpness.

cs.LG↗

CMT-Aware Channel Modeling and Transmit-Power Minimization for Pinching-Antenna Systems

This letter investigates transmit-power minimization for multiuser pinching-antenna system (PAS) from a coupled-mode-theory (CMT)-aware perspective. Existing CMT-based pinching antenna (PA) studies reveal coupling-induced power exchange and radiation behavior, but these effects have not been fully embedded into system-level multi-PA channel modeling and beamforming design. We therefore develop a directional and loss-aware channel model that captures coupling-length-dependent power extraction and the downstream guided-power reduction caused by in-waveguide attenuation and upstream extraction. The model shows that PA design should account for both directional radiation and guided-power evolution, rather than only propagation distance or maximum coupling considered in most existing works. Based on this channel model, we formulate a quality-of-service (QoS)-constrained power minimization problem for continuous PA positioning and finite-codebook activation. For each candidate coupling length, the element-wise positioning and BPSO-based activation use a closed-form zero-forcing (ZF) power metric for low-complexity configuration ranking, thereby avoiding repeated beamforming optimization while excluding rank-deficient candidates and ordering the remaining ones. The selected configuration for each coupling length is then evaluated by optimal fixed-configuration QoS beamforming. Simulation results demonstrate that CMT-aware modeling fundamentally reshapes the preferred PA configuration, maximum coupling is not always power-efficient due to suppressed downstream PA contributions, and finite-codebook activation combined with ZF-based ranking provides a balance between transmit-power performance and deployment complexity.

eess.SP↗

Over-the-Air Interference Nulling Using Active RIS

Interference fundamentally limits the performance of dense wireless networks, and reconfigurable intelligent surfaces (RIS) have recently emerged as a promising means of enabling interference-free transmission in the Degrees-of-Freedom (DoF) sense. This paper investigates the feasibility of achieving full DoF in a two-way K-user interference channel-a canonical interference-limited setting-by employing an active RIS. Unlike its passive counterpart, an active RIS is subject to both per-element gain constraints and a total reflection-power constraint, which renders over-the-air interference nulling equivalent to solving a constrained random linear system with coupled nonlinear constraints. By leveraging tools from high-dimensional convex geometry, we derive a tight scaling threshold on the required number of reflecting elements (REs) for full-DoF transmission. We further extend the analysis to scenarios where each RE incurs circuit power consumption under a total power budget, leading to a fundamental tradeoff between RIS transmit power and circuit power. For this setting, we establish the thresholds for both the total power and the corresponding number of REs required to achieve interference-free transmission. Simulation results validate the theoretical analysis.

cs.IT↗

Over-the-Air Interference Nulling Using Passive RIS for Two-Way K-User Interference Channel

Interference constitutes the fundamental performance bottleneck in wireless networks. Meanwhile, reconfigurable intelligent surface (RIS) has emerged as a promising technique for interference mitigation by directly modifying wireless channels. In this paper, we are interested in the following problem: whether \textit{interference-free} transmission (in terms of Degree-of-Freedom, DoF) can be achieved with the aid of passive RIS in the two-way K-user interference channel, which is regarded as the most severely interfered network. We show that the answer is affirmative, i.e., interference in this network can be neutralized over the air. To accomplish this goal, two prominent challenges arise: i) the unit-modulus constraint on each RIS reflecting coefficient; ii) the significant disparity between the strengths of the direct and reflective channels. To address these challenges, we exploit the high-dimensional and random nature of wireless channels. Specifically, we cast the problem within a high-dimensional convex geometric framework, which enables us to leverage the ubiquitous \textit{concentration} phenomenon in high-dimensional spaces. Based on this framework, we establish both sufficient and necessary conditions on the required number of RIS elements to achieve interference-free DoF, which turns out to \textit{coincide} in order sense. Furthermore, we characterize the impact of imperfect channel state information (CSI) on the achievable DoF and show that interference-free DoF remains achievable if the CSI error is below a certain threshold. Simulation results validate our theoretical findings.

cs.IT↗

How Sparse Can We Prune A Deep Network: A Fundamental Limit Perspective

Network pruning is a commonly used measure to alleviate the storage and computational burden of deep neural networks. However, the fundamental limit of network pruning is still lacking. To close the gap, in this work we'll take a first-principles approach, i.e. we'll directly impose the sparsity constraint on the loss function and leverage the framework of statistical dimension in convex geometry, thus enabling us to characterize the sharp phase transition point, which can be regarded as the fundamental limit of the pruning ratio. Through this limit, we're able to identify two key factors that determine the pruning ratio limit, namely, weight magnitude and network sharpness. Generally speaking, the flatter the loss landscape or the smaller the weight magnitude, the smaller pruning ratio. Moreover, we provide efficient countermeasures to address the challenges in the computation of the pruning limit, which mainly involves the accurate spectrum estimation of a large-scale and non-positive Hessian matrix. Moreover, through the lens of the pruning ratio threshold, we can also provide rigorous interpretations on several heuristics in existing pruning algorithms. Extensive experiments are performed which demonstrate that our theoretical pruning ratio threshold coincides very well with the experiments. All codes are available at: https://github.com/QiaozheZhang/Global-One-shot-Pruning

stat.ML↗

Two-Timescale Learning for Pilot-Free ISAC Systems

A pilot-free integrated sensing and communication (ISAC) system is investigated, in which phase-modulated continuous wave (PMCW) and non-orthogonal multiple access (NOMA) waveforms are co-designed to achieve simultaneous target sensing and data transmission. To enhance effective data throughput (i.e., Goodput) in PMCW-NOMA ISAC systems, we propose a deep learning-based receiver architecture, termed two-timescale Transformer (T3former), which leverages a Transformer architecture to perform joint channel estimation and multi-user signal detection without the need for dedicated pilot signals. By treating the deterministic structure of the PMCW waveform as an implicit pilot, the proposed T3former eliminates the overhead associated with traditional pilot-based methods. The proposed T3former processes the received PMCW-NOMA signals on two distinct timescales, where a fine-grained attention mechanism captures local features across the fast-time dimension, while a coarse-grained mechanism aggregates global spatio-temporal dependencies of the slow-time dimension. Numerical results demonstrate that the proposed T3former significantly outperforms traditional successive interference cancellation (SIC) receivers, which avoids inherent error propagation in SIC. Specifically, the proposed T3former achieves a substantially lower bit error rate and a higher Goodput, approaching the theoretical maximum capacity of a pilot-free system.

cs.IT↗

Analog Self-Interference Cancellation in Full-Duplex Radios: A Fundamental Limit Perspective

Analog self-interference cancellation (A-SIC) plays a crucial role in the implementation of in-band full-duplex (IBFD) radios, due to the fact that the inherent transmit (Tx) noise can only be addressed in the analog domain. It is thus natural to ask what the performance limit of A-SIC is in practical systems, which is still quite underexplored so far. In this paper, we aim to close this gap by characterizing the fundamental performance of A-SIC which employs the common multi-tap delay (MTD) architecture, by accounting for the following practical issues: 1) Nonstationarity of the Tx signal; 2) Nonlinear distortions on the Tx signal; 3) Multipath channel corresponding to the self-interference (SI); 4) Maximum amplitude constraint on the MTD tap weights. Our findings include: 1) The average approximation error for the cyclostationary Tx signals is equal to that for the stationary white Gaussian process, thus greatly simplifying the performance analysis and the optimization procedure. 2) The approximation error for the multipath SI channel can be decomposed as the sum of the approximation error for the single-path scenario. By leveraging these structural results, the optimization framework and algorithms which characterize the fundamental limit of A-SIC, by taking into account all the aforementioned practical factors, are provided.

eess.SP↗

Achieving Interference-Free Degrees of Freedom in Cellular Networks via RIS

It's widely perceived that Reconfigurable Intelligent Surfaces (RIS) cannot increase Degrees of Freedom (DoF) due to their relay nature. A notable exception is Jiang \& Yu's work. They demonstrate via simulation that in an ideal $K$-user interference channel, passive RIS can achieve the interference-free DoF. In this paper, we investigate the DoF gain of RIS in more realistic systems, namely cellular networks, and more challenging scenarios with direct links. We prove that RIS can boost the DoF per cell to that of the interference-free scenario even \textit{ with direct-links}. Furthermore, we \textit{theoretically} quantify the number of RIS elements required to achieve that goal, i.e. $max\left\{ {2L, (\sqrt L + c)η+L } \right\}$ (where $L=GM(GM-1)$, $c$ is a constant and $η$ denotes the ratio of channel strength) for the $G$-cells with more single-antenna users $K$ than base station antennas $M$ per cell. The main challenge lies in addressing the feasibility of a system of algebraic equations, which is difficult by itself in algebraic geometry. We tackle this problem in a probabilistic way, by exploiting the randomness of the involved coefficients and addressing the problem from the perspective of extreme value statistics and convex geometry. Moreover, numerical results confirm the tightness of our theoretical results.

cs.IT↗

Multi-user passive beamforming in RIS-aided communications and experimental validations

Reconfigurable intelligent surface (RIS) is a promising technology for future wireless communications due to its capability of optimizing the propagation environments. Nevertheless, in literature, there are few prototypes serving multiple users. In this paper, we propose a whole flow of channel estimation and beamforming design for RIS, and set up an RIS-aided multi-user system for experimental validations. Specifically, we combine a channel sparsification step with generalized approximate message passing (GAMP) algorithm, and propose to generate the measurement matrix as Rademacher distribution to obtain the channel state information (CSI). To generate the reflection coefficients with the aim of maximizing the spectral efficiency, we propose a quadratic transform-based low-rank multi-user beamforming (QTLM) algorithm. Our proposed algorithms exploit the sparsity and low-rank properties of the channel, which has the advantages of light calculation and fast convergence. Based on the universal software radio peripheral devices, we built a complete testbed working at 5.8GHz and implemented all the proposed algorithms to verify the possibility of RIS assisting multi-user systems. Experimental results show that the system has obtained an average spectral efficiency increase of 13.48bps/Hz, with respective received power gains of 26.6dB and 17.5dB for two users, compared with the case when RIS is powered-off.

eess.SP↗

Multi-level Multiple Instance Learning with Transformer for Whole Slide Image Classification

Whole slide image (WSI) refers to a type of high-resolution scanned tissue image, which is extensively employed in computer-assisted diagnosis (CAD). The extremely high resolution and limited availability of region-level annotations make employing deep learning methods for WSI-based digital diagnosis challenging. Recently integrating multiple instance learning (MIL) and Transformer for WSI analysis shows very promising results. However, designing effective Transformers for this weakly-supervised high-resolution image analysis is an underexplored yet important problem. In this paper, we propose a Multi-level MIL (MMIL) scheme by introducing a hierarchical structure to MIL, which enables efficient handling of MIL tasks involving a large number of instances. Based on MMIL, we instantiated MMIL-Transformer, an efficient Transformer model with windowed exact self-attention for large-scale MIL tasks. To validate its effectiveness, we conducted a set of experiments on WSI classification tasks, where MMIL-Transformer demonstrate superior performance compared to existing state-of-the-art methods, i.e., 96.80% test AUC and 97.67% test accuracy on the CAMELYON16 dataset, 99.04% test AUC and 94.37% test accuracy on the TCGA-NSCLC dataset, respectively. All code and pre-trained models are available at: https://github.com/hustvl/MMIL-Transformer

cs.CV↗

Addressing the curse of mobility in massive MIMO with Prony-based angular-delay domain channel predictions

Massive MIMO is widely touted as an enabling technology for 5th generation (5G) mobile communications and beyond. On paper, the large excess of base station (BS) antennas promises unprecedented spectral efficiency gains. Unfortunately, during the initial phase of industrial testing, a practical challenge arose which threatens to undermine the actual deployment of massive MIMO: user mobility-induced channel Doppler. In fact, testing teams reported that in moderate-mobility scenarios, e.g., 30 km/h of user equipment (UE) speed, the performance drops up to 50% compared to the low-mobility scenario, a problem rooted in the acute sensitivity of massive MIMO to this channel Doppler, and not foreseen by many theoretical papers on the subject. In order to deal with this "curse of mobility", we propose a novel form of channel prediction method, named Prony-based angular-delay domain (PAD) prediction, which is built on exploiting the specific angle-delay-Doppler structure of the multipath. In particular, our method relies on the high angular-delay resolution which arises in the context of 5G. Our theoretical analysis shows that when the number of base station antennas and the bandwidth are large, the prediction error of our PAD algorithm converges to zero for any UE velocity level, provided that only two accurate enough previous channel samples are available. Moreover, when the channel samples are inaccurate, we propose to combine the PAD algorithm with a denoising method for channel estimation phase based on the subspace structure and the long-term statistics of the channel observations. Simulation results show that under a realistic channel model of 3GPP in rich scattering environment, our proposed method is able to overcome this challenge and even approaches the performance of stationary scenarios where the channels do not vary at all.

cs.IT↗

The Impact of Physical Channel on Performance of Subspace-Based Channel Estimation in Massive MIMO Systems

A subspace method for channel estimation has been recently proposed [1] for tackling the pilot contamination effect, which is regarded by some researchers as a bottleneck in massive MIMO systems. It was shown in [1] that if the power ratio between the desired signal and interference is kept above a certain value, the received signal spectrum splits into signal and interference eigenvalues, namely, the "pilot contamination" effect can be completely eliminated. However, [1] assumes an independently distributed (i.d.) channel, which is actually not much the case in practice. Considering this, a more sensible finite-dimensional physical channel model (i.e., a finite scattering environment, where signals impinge on the base station (BS) from a finite number of angles of arrival (AoA)) is employed in this paper. Via asymptotic spectral analysis, it is demonstrated that, compared with the i.d. channel, the physical channel imposes a penalty in the form of an increased power ratio between the useful signal and the interference. Furthermore, we demonstrate an interesting "antenna saturation" effect, i.e., when the number of the BS antennas approaches infinity, the performance under the physical channel with P AoAs is limited by and nearly the same as the performance under the i.d. channel with P receive antennas.

cs.IT↗

A Coordinated Approach to Channel Estimation in Large-scale Multiple-antenna Systems

This paper addresses the problem of channel estimation in multi-cell interference-limited cellular networks. We consider systems employing multiple antennas and are interested in both the finite and large-scale antenna number regimes (so-called "massive MIMO"). Such systems deal with the multi-cell interference by way of per-cell beamforming applied at each base station. Channel estimation in such networks, which is known to be hampered by the pilot contamination effect, constitute a major bottleneck for overall performance. We present a novel approach which tackles this problem by enabling a low-rate coordination between cells during the channel estimation phase itself. The coordination makes use of the additional second-order statistical information about the user channels, which are shown to offer a powerful way of discriminating across interfering users with even strongly correlated pilot sequences. Importantly, we demonstrate analytically that in the large-number-of-antennas regime, the pilot contamination effect is made to vanish completely under certain conditions on the channel covariance. Gains over the conventional channel estimation framework are confirmed by our simulations for even small antenna array sizes.

cs.IT↗