SearcharxivSearch

arXiv subjects

Jing Yee Tan

Publications and source records attributed to Jing Yee Tan.

2 recordsLinked to original sources

Information-Theoretic Bounds for Sparse Covariance Estimation in the Vertical-Split Distributed Model

We study the minimax estimation error for distributed covariance matrix estimation in the vertical-split (feature-split) setting, where two agents each observe different coordinates of~$m$ i.i.d.\ sub-Gaussian samples and communicate a limited number of bits to a central server. While \cite{rahmani2025fundamental} established nearly tight bounds for dense (unstructured) cross-covariance matrices, we investigate whether imposing elementwise $s$-sparsity on the cross-covariance $C_{21}$ can reduce the required communication and sample complexity. In contrast to the horizontal-split setting, where \cite{braverman2016communication} showed that sparsity does \emph{not} reduce communication cost for mean estimation, we prove that sparsity \emph{does} help for cross-covariance estimation in the vertical split. Specifically, for sufficiently large $d_1d_2/s'$ and $0<\varepsilon<σ^2\sqrt{s'}/32$, any scheme achieving expected Frobenius distortion at most $\varepsilon$ must satisfy $B_k = Ω(σ^4 d_k\, s' \log(d_1 d_2/s')/\varepsilon^2)$ and $m = Ω(σ^4\, s' \log(d_1 d_2/s')/\varepsilon^2)$ for cross-covariance estimation, where $s' = s \wedge d_{\min}$. For the $1$-sparse case, our achievable scheme reduces the $d_1d_2$ factor in the dense communication rate to $\log(d_1d_2)$, up to polylogarithmic factors, for the cross-covariance communication component in the matching regime. Our lower bounds are established via Fano's method with an explicit sparse packing using a Varshamov--Gilbert-type argument for signed partial permutation matrices combined with the Conditional Strong Data Processing Inequality of \cite{rahmani2025fundamental}. We show that the communication lower bound is tight up to polylogarithmic factors under the conditions of Remark~\ref{rem:achievmatch}, using an achievable scheme based on covering-net quantization and entry-wise hard thresholding.

cs.IT

Sharp Rates and a One-Line Correction for Spectral Representation Learning

A self-supervised encoder is trained once, frozen, and reused through lightweight probes on tasks nobody named at training time; the practitioner's question is when the off-the-shelf features are good enough and when they need fixing. Canonical correlation analysis, HGR maximal correlation, and the population optimum of the spectral contrastive loss all return the top-$k$ singular subspace of a cross-view dependence operator, justified by isotropy: if the task prior has no directional preference, that subspace is universally optimal. We show isotropy is the wrong hypothesis. The prior enters the transfer risk only through the task covariance $Λ=\mathbb{E}[ΔΔ^\top]$, and only through its compression onto the operator's leading singular directions; what matters is not whether $Λ$ is isotropic but whether its preferred directions are ordered consistently with the operator's spectrum. We prove matching two-sided rates---worst-case regret is exactly $1-1/κ(Λ)$, refines to $1-A_k$ for an alignment coefficient $A_k$, localizes to the top-$2k$ subspace, becomes second order under a spectral gap, and is improvable by no task-agnostic representation---and show why alignment is generic: incoherent preferences cancel in high dimension, and $T$ diverse tasks force $α=\widetilde O(\sqrt{d_x/T})$, a quantitative account of why task diversity, not symmetry, makes self-supervised features transfer. The governing statistics cost $O(kd_x^2)$, and when they signal misalignment a one-line reweighting of the positive-pair term provably restores exact optimality. The result is a diagnostic that answers the practitioner's question from a small labelled budget and refuses when the task bank cannot support the width requested; on controlled data it takes a regret of $0.86$ down to $0.003$, and on a CIFAR-100 encoder it correctly predicts that no correction is needed.

cs.LG