SearcharxivSearch

arXiv subjects

Cosme Louart

Publications and source records attributed to Cosme Louart.

16 recordsLinked to original sources

A Central Limit Theorem for Regularized M-Estimators

We prove a quantitative central limit theorem for linear functionals of regularized empirical-risk minimizers in the proportional-dimensional regime \(p=O(n)\). The data columns are independent, not necessarily identically distributed, and satisfy a uniform columnwise Poincar\'e inequality. Under uniform curvature and smoothness assumptions, and for a quadratic regularizer, we show that every nondegenerate statistic \(\sqrt n\,u^\top\hat\theta\), centered by its expectation and normalized by its standard deviation, converges to a standard normal random variable in Wasserstein distance, with rate \(O((\log n)^7n^{-1/4})\). The proof is based on moment and stability bounds for the minimizer, a second-order leave-one-out expansion, and a perturbative normal-approximation argument for functions of independent variables. We also prove the variance upper bound \(\Var(u^\top\hat\theta)\le C\norm{u}_2^2/n\), identifying the \(\sqrt n\) fluctuation scale.

math.ST

Characterization of Gaussian Universality Breakdown in High-Dimensional Empirical Risk Minimization

We study high-dimensional convex empirical risk minimization (ERM) under general non-Gaussian data designs. By heuristically extending the Convex Gaussian Min-Max Theorem (CGMT) to non-Gaussian settings, we derive an asymptotic min-max characterization of key statistics, enabling approximation of the mean $\mu_{\hat{\theta}}$ and covariance $C_{\hat{\theta}}$ of the ERM estimator $\hat{\theta}$. Specifically, under a concentration assumption on the data matrix and standard regularity conditions on the loss and regularizer, we show that for a test covariate $x$ independent of the training data, the projection $\hat{\theta}^\top x$ approximately follows the convolution of the generally non-Gaussian distribution of $\mu_{\hat{\theta}}^\top x$ with an independent centered Gaussian variable of variance $\mathrm{tr}(C_{\hat{\theta}} \mathbb{E}[xx^\top])$. This result clarifies the scope and limits of Gaussian universality for ERMs. Numerical simulations across diverse losses and models are provided to validate our theoretical predictions and qualitative insights.

stat.ML

Universal concentration for sums under arbitrary dependence

We present a universal concentration bound for sums of random variables under arbitrary dependence, and we prove that it is asymptotically optimal for broad families of marginals admitting a uniform integrable tail-quantile envelope. The bound follows directly from the subadditivity of expected shortfall, a property well known in the risk-measure literature. Our sharpness result relies on an explicit construction of asymptotically extremal couplings. We furthermore provide practical sufficient conditions -- based on convex transformation order comparisons with exponential and power-law envelopes -- under which the bound admits simple, explicit tail profiles.

math.PR

Generalization in Representation Models via Random Matrix Theory: Application to Recurrent Networks

We first study the generalization error of models that use a fixed feature representation (frozen intermediate layers) followed by a trainable readout layer. This setting encompasses a range of architectures, from deep random-feature models to echo-state networks (ESNs) with recurrent dynamics. Working in the high-dimensional regime, we apply Random Matrix Theory to derive a closed-form expression for the asymptotic generalization error. We then apply this analysis to recurrent representations and obtain concise formula that characterize their performance. Surprisingly, we show that a linear ESN is equivalent to ridge regression with an exponentially time-weighted (''memory'') input covariance, revealing a clear inductive bias toward recent inputs. Experiments match predictions: ESNs win in low-sample, short-memory regimes, while ridge prevails with more data or long-range dependencies. Our methodology provides a general framework for analyzing overparameterized models and offers insights into the behavior of deep learning networks.

math.ST

A Random Matrix Perspective of Echo State Networks: From Precise Bias--Variance Characterization to Optimal Regularization

We present a rigorous asymptotic analysis of Echo State Networks (ESNs) in a teacher student setting with a linear teacher with oracle weights. Leveraging random matrix theory, we derive closed form expressions for the asymptotic bias, variance, and mean-squared error (MSE) as functions of the input statistics, the oracle vector, and the ridge regularization parameter. The analysis reveals two key departures from classical ridge regression: (i) ESNs do not exhibit double descent, and (ii) ESNs attain lower MSE when both the number of training samples and the teacher memory length are limited. We further provide an explicit formula for the optimal regularization in the identity input covariance case, and propose an efficient numerical scheme to compute the optimum in the general case. Together, these results offer interpretable theory and practical guidelines for tuning ESNs, helping reconcile recent empirical observations with provable performance guarantees

stat.ML

High-Dimensional Analysis of Bootstrap Ensemble Classifiers

Bootstrap methods have long been the cornerstone of ensemble learning in machine learning. This paper presents a theoretical analysis of bootstrap techniques applied to the Least Square Support Vector Machine (LSSVM) ensemble in the context of large and growing sample sizes and feature dimensionalities. Using tools from Random Matrix Theory, we investigate the performance of this classifier that aggregates decision functions from multiple weak classifiers, each trained on different subsets of the data. We provide insights into the use of bootstrap methods in high-dimensional settings, enhancing our understanding of their impact. Based on these findings, we propose strategies to select the number of subsets and the regularization parameter that maximize the performance of the LSSVM. Empirical experiments on synthetic and real-world datasets validate our theoretical results.

stat.ML

Analysing Multi-Task Regression via Random Matrix Theory with Application to Time Series Forecasting

In this paper, we introduce a novel theoretical framework for multi-task regression, applying random matrix theory to provide precise performance estimations, under high-dimensional, non-Gaussian data distributions. We formulate a multi-task optimization problem as a regularization technique to enable single-task models to leverage multi-task learning information. We derive a closed-form solution for multi-task optimization in the context of linear models. Our analysis provides valuable insights by linking the multi-task learning performance to various model statistics such as raw data covariances, signal-generating hyperplanes, noise levels, as well as the size and number of datasets. We finally propose a consistent estimation of training and testing errors, thereby offering a robust foundation for hyperparameter optimization in multi-task regression scenarios. Experimental validations on both synthetic and real-world datasets in regression and multivariate time series forecasting demonstrate improvements on univariate models, incorporating our method into the training loss and thus leveraging multivariate information.

stat.ML

Operation with Concentration Inequalities

Following the concentration of the measure theory formalism, we consider the transformation $\Phi(Z)$ of a random variable $Z$ having a general concentration function $\alpha$. If the transformation $\Phi$ is $\lambda$-Lipschitz with $\lambda>0$ deterministic, the concentration function of $\Phi(Z)$ is immediately deduced to be equal to $\alpha(\cdot/\lambda)$. If the variations of $\Phi$ are bounded by a random variable $\Lambda$ having a concentration function (around $0$) $\beta: \mathbb R_+\to \mathbb R$, this paper sets that $\Phi(Z)$ has a concentration function analogous to the so-called parallel product of $\alpha$ and $\beta$. With this result at hand (i) we express the concentration of random vectors with independent heavy-tailed entries, (ii) given a transformation $\Phi$ with bounded $k^{\text{th}}$ differential, we express the so-called ``multilevel'' concentration of $\Phi(Z)$ as a function of $\alpha$, and the operator norms of the successive differentials up to the $k^{\text{th}}$ (iii) we obtain a heavy-tailed version of the Hanson--Wright inequality. Finally, in order to rigorously handle the algebraic operations that arise on concentration functions (parallel sums, parallel products, and non-unique pseudo-inverses), we develop at the beginning of the paper a functional framework based on maximally monotone set-valued operators, which provides a natural and coherent formalism for studying these transformations.

math.PR

Concentration of measure and generalized product of random vectors with an application to Hanson-Wright-like inequalities

Starting from concentration of measure hypotheses on $m$ random vectors $Z_1,\ldots, Z_m$, this article provides an expression of the concentration of functionals $ϕ(Z_1,\ldots, Z_m)$ where the variations of $ϕ$ on each variable depend on the product of the norms (or semi-norms) of the other variables (as if $ϕ$ were a product). We illustrate the importance of this result through various generalizations of the Hanson-Wright concentration inequality as well as through a study of the random matrix $XDX^T$ and its resolvent $Q = (I_p - \frac{1}{n}XDX^T)^{-1}$, where $X$ and $D$ are random, which have fundamental interest in statistical machine learning applications.

math.PR

A Concentration of Measure Framework to study convex problems and other implicit formulation problems in machine learning

This paper provides a framework to show the concentration of solutions $Y^*$ to convex minimizing problem where the objective function $ϕ(X)(Y)$ depends on some random vector $X$ satisfying concentration of measure hypotheses. More precisely, the convex problem translates into a contractive fixed point equation that ensure the transmission of the concentration from $X$ to $Y^*$. This result is of central interest to characterize many machine learning algorithms which are defined through implicit equations (e.g., logistic regression, lasso, boosting, etc.). Based on our framework, we provide precise estimations for the first moments of the solution $Y^*$, when $X= (x_1,\ldots, x_n)$ is a data matrix of independent columns and $ϕ(X)(y)$ writes as a sum $\frac{1}{n}\sum_{i=1}^n h_i(x_i^TY)$. That allows to describe the behavior and performance (e.g., generalization error) of a wide variety of machine learning classifiers.

math.PR

A Concentration of Measure and Random Matrix Approach to Large Dimensional Robust Statistics

This article studies the \emph{robust covariance matrix estimation} of a data collection $X = (x_1,\ldots,x_n)$ with $x_i = \sqrt τ_i z_i + m$, where $z_i \in \mathbb R^p$ is a \textit{concentrated vector} (e.g., an elliptical random vector), $m\in \mathbb R^p$ a deterministic signal and $τ_i\in \mathbb R$ a scalar perturbation of possibly large amplitude, under the assumption where both $n$ and $p$ are large. This estimator is defined as the fixed point of a function which we show is contracting for a so-called \textit{stable semi-metric}. We exploit this semi-metric along with concentration of measure arguments to prove the existence and uniqueness of the robust estimator as well as evaluate its limiting spectral distribution.

math.PR

Sharp Bounds for the Concentration of the Resolvent in Convex Concentration Settings

Considering random matrix $X \in \mathcal M_{p,n}$ with independent columns satisfying the convex concentration properties issued from a famous theorem of Talagrand, we express the linear concentration of the resolvent $Q = (I_p - \frac{1}{n}XX^T) ^{-1}$ around a classical deterministic equivalent with a good observable diameter for the nuclear norm. The general proof relies on a decomposition of the resolvent as a series of powers of $X$.

math.PR

Resolvent convergence for sample second-moment matrices with heterogeneous profiles under quadratic-form control

We study the resolvent \(G^z=\left(\frac{1}{n}XX^{\top}-zI_p\right)^{-1}\), where \(z\in\mathbb{C}\) satisfies \(\Im(z)>0\) and \(X=(x_1,\ldots,x_n)\in\mathbb{R}^{p\times n}\) is a random matrix with independent, but not necessarily identically distributed, columns. The columns are real-valued and have finite second moments, but need not be centered. We identify a deterministic equivalent \(\tilde G^z\) through a finite-dimensional fixed-point system depending on the full second-moment profile \((\mathbb{E}[x_ix_i^{\top}])_{i\in[n]}\). Our quantitative, dimension-dependent bounds are expressed in terms of moments of the centered quadratic forms \(q_i(A):=x_i^{\top}Ax_i-\mathbb{E}[x_i^{\top}Ax_i]\), normalized either by the Hilbert--Schmidt norm or by the operator norm of \(A\). In particular, no independence between the entries of a given column is required. We prove quantitative comparison bounds between \(\operatorname{tr}(BG^z)\) and \(\operatorname{tr}(B\tilde G^z)\) in several regimes. We first consider heterogeneous profiles with uniformly bounded operator norm and bounded aspect ratio. Sharper, profile-adapted arguments then cover uniformly bounded Hilbert--Schmidt second moments without any aspect-ratio condition, a common profile for which additional operator-norm estimates are obtained, and profiles taking \(k\) pairwise commuting values, for which the Hilbert--Schmidt bounds incur a linear loss in \(k\). All probabilistic estimates are global for a fixed \(z\in\mathbb{H}\); the regime \(\Im(z)\downarrow0\) is not considered. The Hilbert--Schmidt estimate yields an explicit random-to-deterministic convergence rate and recovers the Marchenko--Pastur limit in the centered i.i.d. setting under the existence of a moment strictly larger than two.

math.PR

Concentration of Measure and Large Random Matrices with an application to Sample Covariance Matrices

The present work provides an original framework for random matrix analysis based on revisiting the concentration of measure theory from a probabilistic point of view. By providing various notions of vector concentration ($q$-exponential, linear, Lipschitz, convex), a set of elementary tools is laid out that allows for the immediate extension of classical results from random matrix theory involving random concentrated vectors in place of vectors with independent entries. These findings are exemplified here in the context of sample covariance matrices but find a large range of applications in statistical learning and beyond, thanks to the broad adaptability of our hypotheses.

math.PR

Random Matrix Theory Proves that Deep Learning Representations of GAN-data Behave as Gaussian Mixtures

This paper shows that deep learning (DL) representations of data produced by generative adversarial nets (GANs) are random vectors which fall within the class of so-called \textit{concentrated} random vectors. Further exploiting the fact that Gram matrices, of the type $G = X^T X$ with $X=[x_1,\ldots,x_n]\in \mathbb{R}^{p\times n}$ and $x_i$ independent concentrated random vectors from a mixture model, behave asymptotically (as $n,p\to \infty$) as if the $x_i$ were drawn from a Gaussian mixture, suggests that DL representations of GAN-data can be fully described by their first two statistical moments for a wide range of standard classifiers. Our theoretical findings are validated by generating images with the BigGAN model and across different popular deep representation networks.

cs.LG

A Random Matrix Approach to Neural Networks

This article studies the Gram random matrix model $G=\frac1TΣ^{\rm T}Σ$, $Σ=σ(WX)$, classically found in the analysis of random feature maps and random neural networks, where $X=[x_1,\ldots,x_T]\in{\mathbb R}^{p\times T}$ is a (data) matrix of bounded norm, $W\in{\mathbb R}^{n\times p}$ is a matrix of independent zero-mean unit variance entries, and $σ:{\mathbb R}\to{\mathbb R}$ is a Lipschitz continuous (activation) function --- $σ(WX)$ being understood entry-wise. By means of a key concentration of measure lemma arising from non-asymptotic random matrix arguments, we prove that, as $n,p,T$ grow large at the same rate, the resolvent $Q=(G+γI_T)^{-1}$, for $γ>0$, has a similar behavior as that met in sample covariance matrix models, involving notably the moment $Φ=\frac{T}n{\mathbb E}[G]$, which provides in passing a deterministic equivalent for the empirical spectral measure of $G$. Application-wise, this result enables the estimation of the asymptotic performance of single-layer random neural networks. This in turn provides practical insights into the underlying mechanisms into play in random neural networks, entailing several unexpected consequences, as well as a fast practical means to tune the network hyperparameters.

math.PR