SearcharxivSearch

arXiv subjects

Stephane Chretien

Publications and source records attributed to Stephane Chretien.

At least 19 recordsLinked to original sources

A finite sample analysis of the benign overfitting phenomenon for ridge function estimation

Recent extensive numerical experiments in high scale machine learning have allowed to uncover a quite counterintuitive phase transition, as a function of the ratio between the sample size and the number of parameters in the model. As the number of parameters $p$ approaches the sample size $n$, the generalisation error increases, but surprisingly, it starts decreasing again past the threshold $p=n$. This phenomenon, brought to the theoretical community attention in \cite{belkin2019reconciling}, has been thoroughly investigated lately, more specifically for simpler models than deep neural networks, such as the linear model when the parameter is taken to be the minimum norm solution to the least-squares problem, firstly in the asymptotic regime when $p$ and $n$ tend to infinity, see e.g. \cite{hastie2019surprises}, and recently in the finite dimensional regime and more specifically for linear models \cite{bartlett2020benign}, \cite{tsigler2020benign}, \cite{lecue2022geometrical}. In the present paper, we propose a finite sample analysis of non-linear models of \textit{ridge} type, where we investigate the \textit{overparametrised regime} of the double descent phenomenon for both the \textit{estimation problem} and the \textit{prediction} problem. Our results provide a precise analysis of the distance of the best estimator from the true parameter as well as a generalisation bound which complements recent works of \cite{bartlett2020benign} and \cite{chinot2020benign}. Our analysis is based on tools closely related to the continuous Newton method \cite{neuberger2007continuous} and a refined quantitative analysis of the performance in prediction of the minimum $\ell_2$-norm solution.

stat.ML

The dual approach to non-negative super-resolution: impact on primal reconstruction accuracy

We study the problem of super-resolution, where we recover the locations and weights of non-negative point sources from a few samples of their convolution with a Gaussian kernel. It has been recently shown that exact recovery is possible by minimising the total variation norm of the measure. An alternative practical approach is to solve its dual. In this paper, we study the stability of solutions with respect to the solutions to the dual problem. In particular, we establish a relationship between perturbations in the dual variable and the primal variables around the optimiser. This is achieved by applying a quantitative version of the implicit function theorem in a non-trivial way.

math.OC

Hedging parameter selection for basis pursuit

In Compressed Sensing and high dimensional estimation, signal recovery often relies on sparsity assumptions and estimation is performed via $\ell_1$-penalized least-squares optimization, a.k.a. LASSO. The $\ell_1$ penalisation is usually controlled by a weight, also called "relaxation parameter", denoted by $λ$. It is commonly thought that the practical efficiency of the LASSO for prediction crucially relies on accurate selection of $λ$. In this short note, we propose to consider the hyper-parameter selection problem from a new perspective which combines the Hedge online learning method by Freund and Shapire, with the stochastic Frank-Wolfe method for the LASSO. Using the Hedge algorithm, we show that a our simple selection rule can achieve prediction results comparable to Cross Validation at a potentially much lower computational cost.

stat.CO

Average performance analysis of the stochastic gradient method for online PCA

This paper studies the complexity of the stochastic gradient algorithm for PCA when the data are observed in a streaming setting. We also propose an online approach for selecting the learning rate. Simulation experiments confirm the practical relevance of the plain stochastic gradient approach and that drastic improvements can be achieved by learning the learning rate.

math.ST

Feature selection in weakly coherent matrices

A problem of paramount importance in both pure (Restricted Invertibility problem) and applied mathematics (Feature extraction) is the one of selecting a submatrix of a given matrix, such that this submatrix has its smallest singular value above a specified level. Such problems can be addressed using perturbation analysis. In this paper, we propose a perturbation bound for the smallest singular value of a given matrix after appending a column, under the assumption that its initial coherence is not large, and we use this bound to derive a fast algorithm for feature extraction.

cs.LG

Post-Prognostics Decision for Optimizing the Commitment of Fuel Cell Systems

In a post-prognostics decision context, this paper addresses the problem of maximizing the useful life of a platform composed of several parallel machines under service constraint. Application on multi-stack fuel cell systems is considered. In order to propose a solution to the insufficient durability of fuel cells, the purpose is to define a commitment strategy by determining at each time the contribution of each fuel cell stack to the global output so as to satisfy the demand as long as possible. A relaxed version of the problem is introduced, which makes it potentially solvable for very large instances. Results based on computational experiments illustrate the efficiency of the new approach, based on the Mirror Prox algorithm, when compared with a simple method of successive projections onto the constraint sets associated with the problem.

cs.CE

An elementary approach to the problem of column selection in a rectangular matrix

The problem of extracting a well conditioned submatrix from any rectangular matrix (with normalized columns) has been studied for some time in functional and harmonic analysis; see \cite{BourgainTzafriri:IJM87,Tropp:StudiaMath08,Vershynin:IJM01} for methods using random column selection. More constructive approaches have been proposed recently; see the recent contributions of \cite{SpielmanSrivastava:IJM12,Youssef:IMRN14}. The column selection problem we consider in this paper is concerned with extracting a well conditioned submatrix, i.e. a matrix whose singular values all lie in $[1-ε,1+ε]$. We provide individual lower and upper bounds for each singular value of the extracted matrix at the price of conceding only one log factor in the number of columns, when compared to the Restricted Invertibility Theorem of Bourgain and Tzafriri. Our method is fully constructive and the proof is short and elementary.

math.FA

On the generic uniform uniqueness of the LASSO estimator

The LASSO is a variable subset selection procedure in statistical linear regression based on $\ell_1$ penalization of the least-squares operator. Uniqueness of the LASSO is an important issue, especially for the study of the LASSO path. The goal of the present paper is to provide a generic sufficient condition on the design matrix for the LASSO minimizer to be unique. Unlike previous works on the question of uniqueness, our condition only depends on the design matrix. Our study is based on a general position condition on the design matrix which holds with probability one for most experimental models.

math.ST

A lower bound on the expected optimal value of certain random linear programs and application to shortest paths and reliability

The paper studies the expectation of the inspection time in complex aging systems. Under reasonable assumptions, this problem is reduced to studying the expectation of the length of the shortest path in the directed degradation graph of the systems where the parameters are given by a pool of experts. The expectation itself being sometimes out of reach, in closed form or even through Monte Carlo simulations in the case of large systems, we propose an easily computable lower bound. The proposed bound applies to a rather general class of linear programs with random nonnegative costs and is directly inspired from the upper bound of Dyer, Frieze and McDiarmid [Math.Programming {\bf 35} (1986), no.1,3--16].

math.ST

Controllability of complex networks using perturbation theory of extreme singular values

Pinning control on complex dynamical networks has emerged as a very important topic in recent trends of control theory due to the extensive study of collective coupled behaviors and their role in physics, engineering and biology. In practice, real-world networks consists of a large number of vertices and one may only be able to perform a control on a fraction of them only. Controllability of such systems has been addressed in \cite{PorfiriDiBernardo:Automatica08}, where it was reformulated as a global asymptotic stability problem. The goal of this short note is to refine the analysis proposed in \cite{PorfiriDiBernardo:Automatica08} using recent results in singular value perturbation theory.

math.OC

On the restricted invertibility problem with an additional orthogonality constraint for random matrices

The Restricted Invertibility problem is the problem of selecting the largest subset of columns of a given matrix $X$, while keeping the smallest singular value of the extracted submatrix above a certain threshold. In this paper, we address this problem in the simpler case where $X$ is a random matrix but with the additional constraint that the selected columns be almost orthogonal to a given vector $v$. Our main result is a lower bound on the number of columns we can extract from a normalized i.i.d. Gaussian matrix for the worst $v$.

math.PR

On the modified Basis Pursuit reconstruction for Compressed Sensing with partially known support

The goal of this short note is to present a refined analysis of the modified Basis Pursuit ($\ell_1$-minimization) approach to signal recovery in Compressed Sensing with partially known support, as introduced by Vaswani and Lu. The problem is to recover a signal $x \in \mathbb R^p$ using an observation vector $y=Ax$, where $A \in \mathbb R^{n\times p}$ and in the highly underdetermined setting $n\ll p$. Based on an initial and possibly erroneous guess $T$ of the signal's support ${\rm supp}(x)$, the Modified Basis Pursuit method of Vaswani and Lu consists of minimizing the $\ell_1$ norm of the estimate over the indices indexed by $T^c$ only. We prove exact recovery essentially under a Restricted Isometry Property assumption of order 2 times the cardinal of $T^c \cap {\rm supp}(x)$, i.e. the number of missed components.

math.ST

Convex recovery of tensors using nuclear norm penalization

The subdifferential of convex functions of the singular spectrum of real matrices has been widely studied in matrix analysis, optimization and automatic control theory. Convex analysis and optimization over spaces of tensors is now gaining much interest due to its potential applications to signal processing, statistics and engineering. The goal of this paper is to present an applications to the problem of low rank tensor recovery based on linear random measurement by extending the results of Tropp to the tensors setting.

stat.ML

Perturbation bounds on the extremal singular values of a matrix after appending a column

In this paper, we study the perturbation of the extreme singular values of a matrix in the particular case where it is obtained after appending an arbitrary column vector. Such results have many applications in bifurcation theory, signal processing, control theory and many other fields. In the first part of this paper, we review and compare various bounds from recent research papers on this subject. We also present a new lower bound and a new upper bound on the perturbation of the operator norm is provided. Simple proofs are provided, based on the study of the characteristic polynomial rather than on variational methods, as e.g. in \cite{Li-Li}. In a second part of the paper, we present applications to signal processing and control theory.

math.SP

Estimation of Gaussian mixtures in small sample studies using $l_1$ penalization

Many experiments in medicine and ecology can be conveniently modeled by finite Gaussian mixtures but face the problem of dealing with small data sets. We propose a robust version of the estimator based on self-regression and sparsity promoting penalization in order to estimate the components of Gaussian mixtures in such contexts. A space alternating version of the penalized EM algorithm is obtained and we prove that its cluster points satisfy the Karush-Kuhn-Tucker conditions. Monte Carlo experiments are presented in order to compare the results obtained by our method and by standard maximum likelihood estimation. In particular, our estimator is seen to perform better than the maximum likelihood estimator.

stat.CO

On prediction with the LASSO when the design is not incoherent

The LASSO estimator is an $\ell_1$-norm penalized least-squares estimator, which was introduced for variable selection in the linear model. When the design matrix satisfies, e.g. the Restricted Isometry Property, or has a small coherence index, the LASSO estimator has been proved to recover, with high probability, the support and sign pattern of sufficiently sparse regression vectors. Under similar assumptions, the LASSO satisfies adaptive prediction bounds in various norms. The present note provides a prediction bound based on a new index for measuring how favorable is a design matrix for the LASSO estimator. We study the behavior of our new index for matrices with independent random columns uniformly drawn on the unit sphere. Using the simple trick of appending such a random matrix (with the right number of columns) to a given design matrix, we show that a prediction bound similar to \cite[Theorem 2.1]{CandesPlan:AnnStat09} holds without any constraint on the design matrix, other than restricted non-singularity.

math.ST

On the spacings between the successive zeros of the Laguerre polynomials

We propose a simple uniform lower bound on the spacings between the successive zeros of the Laguerre polynomials $L_n^{(α)}$ for all $α>-1$. Our bound is sharp regarding the order of dependency on $n$ and $α$ in various ranges. In particular, we recover the orders given in \cite{ahmed} for $α\in (-1,1]$.

math.CA