SearcharxivSearch

arXiv subjects

Zong Shang

Publications and source records attributed to Zong Shang.

6 recordsLinked to original sources

The Geometry of Statistical Feature Learning in Mean-Field Langevin Dynamics

We introduce a geometric formulation of statistical feature learning for supervised regression. Feature learning is defined through a base--fiber decomposition: the base is the feature-side geometry produced by training, and the fiber is the learned feature space where estimation is performed. We prove this property for spherical mean-field Langevin dynamics, viewed as the Wasserstein gradient flow of a negative entropy-regularized empirical risk. In Gaussian multi-index models, the low-temperature stationary distribution concentrates near the hidden indices, forms a multi-spike structure, and yields parameter recovery with high probability, even though negative entropy regularization penalizes concentration. This concentration has a sharp transition at temperature $\lambda\asymp 1$. In Gaussian single-index models, the stationary measure satisfies a concentration property, with parity determining whether it lives on $S_2^{d-1}$ or $\mathbb{RP}^{d-1}$. The induced learned feature space aligns the regression signal and yields rates $d/N$ and $Md/N$, up to logarithmic factors.

math.ST

Sharp convergence rates for Spectral methods via the feature space decomposition method

In this paper, we apply the Feature Space Decomposition (FSD) method developed in [LS24, GLS25, LSSW26, ALSS26] to obtain, under fairly general conditions, matching upper and lower bounds for the population excess risk of spectral methods in linear regression under the squared loss, for every covariance and every signal. This result enables us, for a given linear regression problem, to define a pre-order on the set of spectral methods according to their convergence rates, thereby characterizing which spectral algorithm is superior for that specific problem. Furthermore, this allows us to generalize the saturation effect proposed in inverse problems and to provide necessary and sufficient conditions for its occurrence. Our method also shows that, under broad conditions, any spectral algorithm cannot overcome the barrier of the information exponent in problems such as single-index learning.

math.ST

Upper bounds for the L^q empirical process via generic chaining

Using the generic chaining method, we derive upper bounds for the \(L^q\) process of sub-Gaussian classes when \(1 \le q \le 2\), thereby resolving an open problem posed by Al-Ghattas, Chen, and Sanz-Alonso in arXiv:2502.16916. Combined with the results of arXiv:2502.16916, this yields upper bounds for the \(L^q\) process for all \(1 \le q < \infty\). We also present corollaries of this result in the geometry of Banach spaces, including high-probability bounds on the \(\ell_q\) norm diameter of random hyperplane sections of convex bodies where the subspaces are not necessarily uniformly distributed on the Grassmannian manifold and the restricted isomorphic property for \(\ell_q\) norm.

math.PR

A Geometrical Analysis of Kernel Ridge Regression and its Applications

We obtain upper bounds for the estimation error of Kernel Ridge Regression (KRR) for all non-negative regularization parameters, offering a geometric perspective on various phenomena in KRR. As applications: 1. We address the multiple descent problem, unifying the proofs of arxiv:1908.10292 and arxiv:1904.12191 for polynomial kernels and we establish multiple descent for the upper bound of estimation error of KRR under sub-Gaussian design and non-asymptotic regimes. 2. For a sub-Gaussian design vector and for non-asymptotic scenario, we prove a one-sided isomorphic version of the Gaussian Equivalent Conjecture. 3. We offer a novel perspective on the linearization of kernel matrices of non-linear kernel, extending it to the power regime for polynomial kernels. 4. Our theory is applicable to data-dependent kernels, providing a convenient and accurate tool for the feature learning regime in deep learning theory. 5. Our theory extends the results in arxiv:2009.14286 under weak moment assumption. Our proof is based on three mathematical tools developed in this paper that can be of independent interest: 1. Dvoretzky-Milman theorem for ellipsoids under (very) weak moment assumptions. 2. Restricted Isomorphic Property in Reproducing Kernel Hilbert Spaces with embedding index conditions. 3. A concentration inequality for finite-degree polynomial kernel functions.

math.ST

A geometrical viewpoint on the benign overfitting property of the minimum $l_2$-norm interpolant estimator and its universality

In the linear regression model, the minimum l2-norm interpolant estimator has received much attention since it was proved to be consistent even though it fits noisy data perfectly under some condition on the covariance matrix $Σ$ of the input vector, known as benign overfitting. Motivated by this phenomenon, we study the generalization property of this estimator from a geometrical viewpoint. Our main results extend and improve the convergence rates as well as the deviation probability from [Tsigler and Bartlett]. Our proof differs from the classical bias/variance analysis and is based on the self-induced regularization property introduced in [Bartlett, Montanari and Rakhlin]: the minimum l2-norm interpolant estimator can be written as a sum of a ridge estimator and an overfitting component. The two geometrical properties of random Gaussian matrices at the heart of our analysis are the Dvoretsky-Milman theorem and isomorphic and restricted isomorphic properties. In particular, the Dvoretsky dimension appearing naturally in our geometrical viewpoint, coincides with the effective rank and is the key tool for handling the behavior of the design matrix restricted to the sub-space where overfitting happens. We extend these results to heavy-tailed scenarii proving the universality of this phenomenon beyond exponential moment assumptions. This phenomenon is unknown before and is widely believed to be a significant challenge. This follows from an anistropic version of the probabilistic Dvoretsky-Milman theorem that holds for heavy-tailed vectors which is of independent interest.

math.ST

Benign overfitting without concentration

We obtain a sufficient condition for benign overfitting of linear regression problem. Our result does not rely on concentration argument but on small-ball assumption and thus can holds in heavy-tailed case. The basic idea is to establish a coordinate small-ball estimate in terms of effective rank so that we can calibrate the balance of epsilon-Net and exponential probability. Our result indicates that benign overfitting is not depending on concentration property of the input vector. Finally, we discuss potential difficulties for benign overfitting beyond linear model and a benign overfitting result without truncated effective rank.

math.ST