SearcharxivSearch

arXiv subjects

Pei-Chun Su

Publications and source records attributed to Pei-Chun Su.

8 recordsLinked to original sources

eOptShrinkQ: Near-Lossless KV Cache Compression Through Optimal Spectral Denoising and Quantization

We show that the key-value (KV) cache in transformer attention heads admits a natural decomposition into a low-rank \emph{shared context} component and a full-rank \emph{per-token} residual, well described by the spiked random matrix model. This observation leads to eOptShrinkQ, a two-stage compression pipeline: optimal singular value shrinkage (eOptShrink) automatically extracts the shared structure, and the residual -- which satisfies the \emph{thin shell property} with delocalized coordinates -- is quantized by TurboQuant~\citep{zandieh2025turboquant}, a recently proposed per-vector scalar quantizer with near-optimal distortion guarantees. By restoring the isotropy that scalar quantization assumes, spectral denoising eliminates the need for both outlier handling and dedicated inner product bias correction, freeing those bits for improved reconstruction. The theoretical grounding in random matrix theory provides three guarantees: automatic rank selection via the BBP phase transition, provably near-zero inner product bias on the residual, and coordinate delocalization ensuring near-optimal quantization distortion. Experimentally, we validate eOptShrinkQ on Llama-3.1-8B and Ministral-8B across three levels: per-head MSE and inner product fidelity, where eOptShrinkQ saves nearly one bit per entry over TurboQuant at equivalent quality; end-to-end on LongBench (16 tasks), where eOptShrinkQ at $\sim$2.2 bits per entry outperforms TurboQuant at 3.0 bits; and multi-needle retrieval, where eOptShrinkQ at 2.2 bits closely matches or exceeds uncompressed FP16, suggesting that spectral denoising can act as a beneficial regularizer for retrieval-intensive tasks.

cs.LG

Data-Driven Matrix Recovery via Optimal Shrinkage and Spatially Resolved Singular Vector Denoising under High-Dimensional Separable Noise

This paper develops a spatially resolved perturbation theory for singular vectors under high-dimensional separable noise and applies it to data-driven matrix recovery. In the asymptotic regime where the matrix dimensions are proportional and significantly larger than the signal rank, we derive exact leading-order variance formulas for the singular vector perturbation projected onto any spatial patch. The variance decomposes into a spatially non-uniform component governed by the local noise covariance and a spatially uniform component governed by the global noise level. These formulas provide the foundation for the \emph{extended optimal shrinkage and wavelet shrinkage} (e$\mathcal{OWS}$) algorithm, which recovers low-rank matrices satisfying a mixed Hölder condition. The pipeline begins with optimal shrinkage of singular values, then constructs coupled multiscale partition trees on the row and column spaces from the denoised estimate, generating a tensor Haar-Walsh wavelet basis. Spatially adaptive wavelet shrinkage is applied using data-driven, coefficient-level thresholds derived from the perturbation theory. We establish convergence rates that strictly improve upon both optimal shrinkage and wavelet shrinkage applied in isolation. Numerical simulations demonstrate reliable matrix recovery and accurate reconstruction of the underlying singular subspaces, including an application to fetal ECG extraction.

math.SP

Extracting Dual Analytic Geometries of Linear Transformations to Achieve Efficient Computation

We propose a novel framework for fast integral operations by uncovering hidden geometries in the row and column structures of the underlying operators. This is accomplished through the \texttt{Questionnaire} algorithm, an iterative procedure that constructs adaptive hierarchical partition trees, revealing latent multiscale organizations and exposing local low-rank structures within the data. Guided by these geometries, we employ two complementary techniques: (1) The \texttt{\texttt{Butterfly}} algorithm, which exploits the learned hierarchical low-rank structure; and (2) Adaptive \texttt{eGHWT}, best tilings in both space and frequency using all levels of the generalized Haar--Walsh wavelet packets. These techniques enable efficient matrix factorization and multiplication. We coin our algorithms as \texttt{Questionnaire Factorization and Fast Transform (QFFT)}. Unlike classical approaches that rely on prior knowledge of the underlying geometry, \texttt{QFFT} is fully data-driven and applicable to matrices arising from irregular or unknown distributions. Even when the rows and columns both appear mutually orthogonal, our framework identifies the intrinsic ordering of orthogonal vectors that reveal hidden sparsity of the kernel. We demonstrate the effectiveness of our approach on matrices associated with heterogeneous operators and families of orthogonal polynomials. The resulting compressed representations reduce storage complexity from $\mathcal{O}(N^2)$ to $\mathcal{O}(N \log N)$, enabling fast computation and scalable implementation.

math.NA

Intrinsic and Extrinsic Organized Attention: Softmax Invariance and Network Sparsity

We examine the intrinsic (within the attention head) and extrinsic (amongst the attention heads) structure of the self-attention mechanism in transformers. Theoretical evidence for invariance of the self-attention mechanism to softmax activation is obtained by appealing to paradifferential calculus, (and is supported by computational examples), which relies on the intrinsic organization of the attention heads. Furthermore, we use an existing methodology for hierarchical organization of tensors to examine network structure by constructing hierarchal partition trees with respect to the query, key, and head axes of network 3-tensors. Such an organization is consequential since it allows one to profitably execute common signal processing tasks on a geometry where the organized network 3-tensors exhibit regularity. We exemplify this qualitatively, by visualizing the hierarchical organization of the tree comprised of attention heads and the diffusion map embeddings, and quantitatively by investigating network sparsity with the expansion coefficients of individual attention heads and the entire network with respect to the bi and tri-haar bases (respectively) on the space of queries, keys, and heads of the network. To showcase the utility of our theoretical and methodological findings, we provide computational examples using vision and language transformers. The ramifications of these findings are two-fold: (1) a subsequent step in interpretability analysis is theoretically admitted, and can be exploited empirically for downstream interpretability tasks (2) one can use the network 3-tensor organization for empirical network applications such as model pruning (by virtue of network sparsity) and network architecture comparison.

math.NA

Data-Driven optimal shrinkage of singular values under high-dimensional noise with separable covariance structure with application

We develop a data-driven optimal shrinkage algorithm for matrix denoising in the presence of high-dimensional noise with a separable covariance structure; that is, the noise is colored and dependent across samples. The algorithm, coined {\em extended OptShrink} (eOptShrink) depends on the asymptotic behavior of singular values and singular vectors of the random matrix associated with the noisy data. Based on the developed theory, including the sticking property of non-outlier singular values and delocalization of the non-outlier singular vectors associated with weak signals with a convergence rate, and the spectral behavior of outlier singular values and vectors, we develop three estimators, each of these has its own interest. First, we design a novel rank estimator, based on which we provide an estimator for the spectral distribution of the pure noise matrix, and hence the optimal shrinker called eOptShrink. In this algorithm we do not need to estimate the separable covariance structure of the noise. A theoretical guarantee of these estimators with a convergence rate is given. On the application side, in addition to a series of numerical simulations with a comparison with various state-of-the-art optimal shrinkage algorithms, we apply eOptShrink to extract maternal and fetal electrocardiograms from the single channel trans-abdominal maternal electrocardiogram.

stat.AP

Optimal Recovery of Precision Matrix for Mahalanobis Distance from High Dimensional Noisy Observations in Manifold Learning

Motivated by establishing theoretical foundations for various manifold learning algorithms, we study the problem of Mahalanobis distance (MD), and the associated precision matrix, estimation from high-dimensional noisy data. By relying on recent transformative results in covariance matrix estimation, we demonstrate the sensitivity of \MD~and the associated precision matrix to measurement noise, determining the exact asymptotic signal-to-noise ratio at which MD fails, and quantifying its performance otherwise. In addition, for an appropriate loss function, we propose an asymptotically optimal shrinker, which is shown to be beneficial over the classical implementation of the MD, both analytically and in simulations. The result is extended to the manifold setup, where the nonlinear interaction between curvature and high-dimensional noise is taken care of. The developed solution is applied to study a multiscale reduction problem in the dynamical system analysis.

math.ST

Recovery of the fetal electrocardiogram for morphological analysis from two trans-abdominal channels via optimal shrinkage

We propose a novel algorithm to recover fetal electrocardiogram (ECG) for both the fetal heart rate analysis and morphological analysis of its waveform from two or three trans-abdominal maternal ECG channels. We design an algorithm based on the optimal-shrinkage and the nonlocal Euclidean median under the wave-shape manifold model. For the fetal heart rate analysis, the algorithm is evaluated on publicly available database, 2013 PhyioNet/Computing in Cardiology Challenge, set A. For the morphological analysis, we propose to simulate semi-real databases by mixing the MIT-BIH Normal Sinus Rhythm Database and MITDB Arrhythmia Database. For the fetal R peak detection, the proposed algorithm outperforms all algorithms under comparison. For the morphological analysis, the algorithm provides an encouraging result in recovery of the fetal ECG waveform, including PR, QT and ST intervals, even when the fetus has arrhythmia. To the best of our knowledge, this is the first work focusing on recovering the fetal ECG for morphological analysis from two or three channels with an algorithm potentially applicable for continuous fetal electrocardiographic monitoring, which creates the potential for long term monitoring purpose.

eess.SP

Fetus: the radar of maternal stress, a cohort study

Objective: We hypothesized that prenatal stress (PS) exerts lasting impact on fetal heart rate (fHR). We sought to validate the presence of such PS signature in fHR by measuring coupling between maternal HR (mHR) and fHR. Study design: Prospective observational cohort study in stressed group (SG) mothers with controls matched for gestational age during screening at third trimester using Cohen Perceived Stress Scale (PSS) questionnaire with PSS-10 equal or above 19 classified as SG. Women with PSS-10 less than 19 served as control group (CG). Setting: Klinikum rechts der Isar of the Technical University of Munich. Population: Singleton 3rd trimester pregnant women. Methods: Transabdominal fetal electrocardiograms (fECG) were recorded. We deployed a signal processing algorithm termed bivariate phase-rectified signal averaging (BPRSA) to quantify coupling between mHR and fHR resulting in a fetal stress index (FSI). Maternal hair cortisol was measured at birth. Differences were assumed to be significant for p value less than 0.05. Main Outcome Measures: Differences for FSI between both groups. Results: We screened 1500 women enrolling 538 of which 16.5 % showed a PSS-10 score equal or above 19 at 34+0 weeks. Fifty five women eventually comprised the SG and n=55 served as CG. Median PSS was 22.0 (IQR 21.0-24.0) in the SG and 9.0 (6.0-12.0) in the CG, respectively. Maternal hair cortisol was higher in SG than CG at 86.6 (48.0-169.2) versus 53.0 (34.4-105.9) pg/mg. At 36+5 weeks, FSI was significantly higher in fetuses of stressed mothers when compared to controls [0.43 (0.18-0.85) versus 0.00 (-0.49-0.18)]. Conclusion: Our findings show a persistent effect of PS affecting fetuses in the last trimester.

q-bio.QM