SearcharxivSearch

arXiv subjects

Yannick Baraud

Publications and source records attributed to Yannick Baraud.

At least 19 recordsLinked to original sources

Statistical Inference via T-Posterior Randomised Estimators

Given a statistical model, we propose a novel estimation method that yields randomised estimators for the unknown distribution of an observed random variable. We establish non-asymptotic bounds for the performance of these estimators and demonstrate their robustness to potential model misspecification. Notably, these properties are established by circumventing the use of concentration inequalities and empirical process theory. We provide an illustration of this approach to the problem of estimating the intensity of a Poisson process.

math.ST

Estimating a regression function under possible heteroscedastic and heavy-tailed errors. Application to shape-restricted regression

We consider a regression framework where the design points are deterministic and the errors possibly non-i.i.d. and heavy-tailed (with a moment of order $p$ in $[1,2]$). Given a class of candidate regression functions, we propose a surrogate for the classical least squares estimator (LSE). For this new estimator, we establish a nonasymptotic risk bound with respect to the absolute loss which takes the form of an oracle type inequality. This inequality shows that our estimator possesses natural adaptation properties with respect to some elements of the class. When this class consists of monotone functions or convex functions on an interval, these adaptation properties are similar to those established in the literature for the LSE. However, unlike the LSE, we prove that our estimator remains stable with respect to a possible heteroscedasticity of the errors and may even converge at a parametric rate (up to a logarithmic factor) when the LSE is not even consistent. We illustrate the performance of this new estimator over classes of regression functions that satisfy a shape constraint: piecewise monotone, piecewise convex/concave, among other examples. The paper also contains some approximation results by splines with degrees in $\{0,1\}$ and VC bounds for the dimensions of classes of level sets. These results may be of independent interest.

math.ST

Robust estimation of a regression function in exponential families

We observe $n$ pairs of independent (but not necessarily i.i.d.) random variables $X_{1}=(W_{1},Y_{1}),\ldots,X_{n}=(W_{n},Y_{n})$ and tackle the problem of estimating the conditional distributions $Q_{i}^{\star}(w_{i})$ of $Y_{i}$ given $W_{i}=w_{i}$ for all $i\in\{1,\ldots,n\}$. Even though these might not be true, we base our estimator on the assumptions that the data are i.i.d.\ and the conditional distributions of $Y_{i}$ given $W_{i}=w_{i}$ belong to a one parameter exponential family $\bar{\mathscr{Q}}$ with parameter space given by an interval $I$. More precisely, we pretend that these conditional distributions take the form $Q_{{\boldsymbolθ}(w_{i})}\in \bar{\mathscr{Q}}$ for some ${\boldsymbolθ}$ that belongs to a VC-class $\bar{\boldsymbolΘ}$ of functions with values in $I$. For each $i\in\{1,\ldots,n\}$, we estimate $Q_{i}^{\star}(w_{i})$ by a distribution of the same form, i.e.\ $Q_{\hat{\boldsymbolθ}(w_{i})}\in \bar{\mathscr{Q}}$, where $\hat {\boldsymbolθ}=\hat {\boldsymbolθ}(X_{1},\ldots,X_{n})$ is a well-chosen estimator with values in $\bar{\boldsymbolΘ}$. We show that our estimation strategy is robust to model misspecification, contamination and the presence of outliers. Besides, we provide an algorithm for calculating $\hat{\boldsymbolθ}$ when $\bar{\boldsymbolΘ}$ is a VC-class of functions of low or moderate dimension and we carry out a simulation study to compare the performance of $\hat{\boldsymbolθ}$ to that of the MLE and median-based estimators.

math.ST

From robust tests to Bayes-like posterior distributions

In the Bayes paradigm and for a given loss function, we propose the construction of a new type of posterior distributions, that extends the classical Bayes one, for estimating the law of an $n$-sample. The loss functions we have in mind are based on the total variation and Hellinger distances as well as some $\mathbb{L}_{j}$-ones. We prove that, with a probability close to one, this new posterior distribution concentrates its mass in a neighbourhood of the law of the data, for the chosen loss function, provided that this law belongs to the support of the prior or, at least, lies close enough to it. We therefore establish that the new posterior distribution enjoys some robustness properties with respect to a possible misspecification of the prior, or more precisely, its support. For the total variation and squared Hellinger losses, we also show that the posterior distribution keeps its concentration properties when the data are only independent, hence not necessarily i.i.d., provided that most of their marginals or the average of these are close enough to some probability distribution around which the prior puts enough mass. The posterior distribution is therefore also stable with respect to the equidistribution assumption. We illustrate these results by several applications. We consider the problems of estimating a location parameter or both the location and the scale of a density in a nonparametric framework. Finally, we also tackle the problem of estimating a density, with the squared Hellinger loss, in a high-dimensional parametric model under some sparsity conditions. The results established in this paper are non-asymptotic and provide, as much as possible, explicit constants.

math.ST

Tests and estimation strategies associated to some loss functions

We consider the problem of estimating the joint distribution of $n$ independent random variables. Our approach is based on a family of candidate probabilities that we shall call a model and which is chosen to either contain the true distribution of the data or at least to provide a good approximation of it with respect to some loss function. The aim of the present paper is to describe a general estimation strategy that allows to adapt to both the specific features of the model and the choice of the loss function in view of designing an estimator with good estimation properties. The losses we have in mind are based on the total variation, Hellinger, Wasserstein and $\mathbb{L}_p$-distances to name a few. We show that the risk of the resulting estimator with respect to the loss function can be bounded by the sum of an approximation term accounting for the loss between the true distribution and the model and a complexity term that corresponds to the bound we would get if this distribution did belong to the model. Our results hold under mild assumptions on the true distribution of the data and are based on exponential deviation inequalities that are non-asymptotic and involve explicit constants. When the model reduces to two distinct probabilities, we show how our estimation strategy leads to a robust test whose errors of first and second kinds only depend on the losses between the true distribution and the two tested probabilities.

math.ST

Robust Bayes-Like Estimation: Rho-Bayes estimation

We consider the problem of estimating the joint distribution $P$ of $n$ independent random variables within the Bayes paradigm from a non-asymptotic point of view. Assuming that $P$ admits some density $s$ with respect to a given reference measure, we consider a density model $\overline S$ for $s$ that we endow with a prior distribution $π$ (with support $\overline S$) and we build a robust alternative to the classical Bayes posterior distribution which possesses similar concentration properties around $s$ whenever it belongs to the model $\overline S$. Furthermore, in density estimation, the Hellinger distance between the classical and the robust posterior distributions tends to 0, as the number of observations tends to infinity, under suitable assumptions on the model and the prior, provided that the model $\overline S$ contains the true density $s$. However, unlike what happens with the classical Bayes posterior distribution, we show that the concentration properties of this new posterior distribution are still preserved in the case of a misspecification of the model, that is when $s$ does not belong to $\overline S$ but is close enough to it with respect to the Hellinger distance.

math.ST

About the lower bounds for the multiple testing problem

Given an observed random variable, consider the problem of recovering its distribution among a family of candidate ones. The two-point inequality, Fano's lemma and more recently an inequality due to Venkataramanan and Johnson (2018) allow to bound the maximal probability of error over the family from below. The aim of this paper is to give a very short and simple proof of all these results simultaneously and improve in passing the inequality of Venkataramanan and Johnson.

math.ST

Rho-estimators revisited: General theory and applications

Following Baraud, Birgé and Sart (2017), we pursue our attempt to design a robust universal estimator of the joint ditribution of $n$ independent (but not necessarily i.i.d.) observations for an Hellinger-type loss. Given such observations with an unknown joint distribution $\mathbf{P}$ and a dominated model $\mathscr{Q}$ for $\mathbf{P}$, we build an estimator $\widehat{\mathbf{P}}$ based on $\mathscr{Q}$ and measure its risk by an Hellinger-type distance. When $\mathbf{P}$ does belong to the model, this risk is bounded by some quantity which relies on the local complexity of the model in a vicinity of $\mathbf{P}$. In most situations this bound corresponds to the minimax risk over the model (up to a possible logarithmic factor). When $\mathbf{P}$ does not belong to the model, its risk involves an additional bias term proportional to the distance between $\mathbf{P}$ and $\mathscr{Q}$, whatever the true distribution $\mathbf{P}$. From this point of view, this new version of $ρ$-estimators improves upon the previous one described in Baraud, Birgé and Sart (2017) which required that $\mathbf{P}$ be absolutely continuous with respect to some known reference measure. Further additional improvements have been brought as compared to the former construction. In particular, it provides a very general treatment of the regression framework with random design as well as a computationally tractable procedure for aggregating estimators. We also give some conditions for the Maximum Likelihood Estimator to be a $ρ$-estimator. Finally, we consider the situation where the Statistician has at disposal many different models and we build a penalized version of the $ρ$-estimator for model selection and adaptation purposes. In the regression setting, this penalized estimator not only allows to estimate the regression function but also the distribution of the errors.

math.ST

Une alternative robuste au maximum de vraisemblance: la $ρ$-estimation

This paper is based on our personal notes for the short course we gave on January 5, 2017 at Institut Henri Poincaré, after an invitation of the SFdS. Our purpose is to give an overview of the method of $ρ$-estimation and of the optimality and robustness properties of the estimators built according to this procedure. This method can be viewed as the sequel of a long series of researches which were devoted to the construction of estimators with good properties in various statistical frameworks. We shall emphasize the connection between the $ρ$-estimators and the previous ones, in particular the maximum likelihood estimator, and we shall show, via some typical examples, that the $ρ$-estimators perform better from various points of view. ------ Cet article est fondé sur les notes du mini-cours que nous avons donné le 5 janvier 2017 à l'Institut Henri Poincaré à l'occasion d'une journée organisée par la SFdS et consacrée à la Statistique Mathématique. Il vise à donner un aperçu de la méthode de $ρ$-estimation ainsi que des propriétés d'optimalité et de robustesse des estimateurs construits selon cette procédure. Cette méthode s'inscrit dans une longue lignée de recherches dont l'objectif a été de produire des estimateurs possédant de bonnes propriétés pour un ensemble de cadres statistiques aussi vaste que possible. Nous mettrons en lumière les liens forts qui existent entre les $ρ$-estimateurs et ces prédécesseurs, notamment les estimateurs du maximum de vraisemblance, mais montrerons également, au travers d'exemples choisis, que les $ρ$-estimateurs les surpassent sur bien des aspects.

math.ST

A new method for estimation and model selection: $ρ$-estimation

The aim of this paper is to present a new estimation procedure that can be applied in many statistical frameworks including density and regression and which leads to both robust and optimal (or nearly optimal) estimators. In density estimation, they asymptotically coincide with the celebrated maximum likelihood estimators at least when the statistical model is regular enough and contains the true density to estimate. For very general models of densities, including non-compact ones, these estimators are robust with respect to the Hellinger distance and converge at optimal rate (up to a possible logarithmic factor) in all cases we know. In the regression setting, our approach improves upon the classical least squares from many aspects. In simple linear regression for example, it provides an estimation of the coefficients that are both robust to outliers and simultaneously rate-optimal (or nearly rate-optimal) for large class of error distributions including Gaussian, Laplace, Cauchy and uniform among others.

math.ST

Rates of convergence of rho-estimators for sets of densities satisfying shape constraints

The purpose of this paper is to pursue our study of rho-estimators built from i.i.d. observations that we defined in Baraud et al. (2014). For a ρ-estimator based on some model S (which means that the estimator belongs to S) and a true distribution of the observations that also belongs to S, the risk (with squared Hellinger loss) is bounded by a quantity which can be viewed as a dimension function of the model and is often related to the "metric dimension" of this model, as defined in Birgé (2006). This is a minimax point of view and it is well-known that it is pessimistic. Typically, the bound is accurate for most points in the model but may be very pessimistic when the true distribution belongs to some specific part of it. This is the situation that we want to investigate here. For some models, like the set of decreasing densities on [0,1], there exist specific points in the model that we shall call "extremal" and for which the risk is substantially smaller than the typical risk. Moreover, the risk at a non-extremal point of the model can be bounded by the sum of the risk bound at a well-chosen extremal point plus the square of its distance to this point. This implies that if the true density is close enough to an extremal point, the risk at this point may be smaller than the minimax risk on the model and this actually remains true even if the true density does not belong to the model. The result is based on some refined bounds on the suprema of empirical processes that are established in Baraud (2016).

math.ST

Bounding the expectation of the supremum of an empirical process over a (weak) vc-major class

Given a bounded class of functions G and independent random variables X1, . . . , Xn, we provide an upper bound for the expectation of the supremum of the empirical process over elements of G having a small variance. Our bound applies in the cases where G is a VC-subgraph or a VC-major class and it is of smaller order than those one could get by using a universal entropy bound over the whole class G . It also involves explicit constants and does not require the knowledge of the entropy of G

math.PR

Estimation of the density of a determinantal process

We consider the problem of estimating the density $Π$ of a determinantal process $N$ from the observation of $n$ independent copies of it. We use an aggregation procedure based on robust testing to build our estimator. We establish non-asymptotic risk bounds with respect to the Hellinger loss and deduce, when $n$ goes to infinity, uniform rates of convergence over classes of densities $Π$ of interest.

math.ST

Estimating composite functions by model selection

We consider the problem of estimating a function $s$ on $[-1,1]^{k}$ for large values of $k$ by looking for some best approximation by composite functions of the form $g\circ u$. Our solution is based on model selection and leads to a very general approach to solve this problem with respect to many different types of functions $g,u$ and statistical frameworks. In particular, we handle the problems of approximating $s$ by additive functions, single and multiple index models, neural networks, mixtures of Gaussian densities (when $s$ is a density) among other examples. We also investigate the situation where $s=g\circ u$ for functions $g$ and $u$ belonging to possibly anisotropic smoothness classes. In this case, our approach leads to a completely adaptive estimator with respect to the regularity of $s$.

math.ST

Estimator selection in the Gaussian setting

We consider the problem of estimating the mean $f$ of a Gaussian vector $Y$ with independent components of common unknown variance $σ^{2}$. Our estimation procedure is based on estimator selection. More precisely, we start with an arbitrary and possibly infinite collection $\FF$ of estimators of $f$ based on $Y$ and, with the same data $Y$, aim at selecting an estimator among $\FF$ with the smallest Euclidean risk. No assumptions on the estimators are made and their dependencies with respect to $Y$ may be unknown. We establish a non-asymptotic risk bound for the selected estimator. As particular cases, our approach allows to handle the problems of aggregation and model selection as well as those of choosing a window and a kernel for estimating a regression function, or tuning the parameter involved in a penalized criterion. We also derive oracle-type inequalities when $\FF$ consists of linear estimators. For illustration, we carry out two simulation studies. One aims at comparing our procedure to cross-validation for choosing a tuning parameter. The other shows how to implement our approach to solve the problem of variable selection in practice.

math.ST

A Bernstein-type inequality for suprema of random processes with applications to model selection in non-Gaussian regression

Let $\pa{X_{t}}_{t\in T}$ be a family of real-valued centered random variables indexed by a countable set $T$. In the first part of this paper, we establish exponential bounds for the deviation probabilities of the supremum $Z=\sup_{t\in T}X_{t}$ by using the generic chaining device introduced in Talagrand (2005). Compared to concentration-type inequalities, these bounds offer the advantage to hold under weaker conditions on the family $\pa{X_{t}}_{t\in T}$. The second part of the paper is oriented towards statistics. We consider the regression setting $Y=f+\eps$ where $f$ is an unknown vector of $\R^{n}$ and $\eps$ is a random vector the components of which are independent, centered and admit finite Laplace transforms in a neighborhood of 0. Our aim is to estimate $f$ from the observation of $Y$ by mean of a model selection approach among a collection of linear subspaces of $\R^{n}$. The selection procedure we propose is based on the minimization of a penalized criterion the penalty of which is calibrated by using the deviation bounds established in the first part of this paper. More precisely, we study suprema of random variables of the form $X_{t}=\sum_{i=1}^{n}t_{i}\eps_{i}$ when $t$ varies among the unit ball of a linear subspace of $\R^{n}$. We finally show that our estimator satisfies some oracle-type inequality under suitable assumptions on the metric structures of the linear spaces of the collection.

math.ST

Estimator selection with respect to Hellinger-type risks

We observe a random measure $N$ and aim at estimating its intensity $s$. This statistical framework allows to deal simultaneously with the problems of estimating a density, the marginals of a multivariate distribution, the mean of a random vector with nonnegative components and the intensity of a Poisson process. Our estimation strategy is based on estimator selection. Given a family of estimators of $s$ based on the observation of $N$, we propose a selection rule, based on $N$ as well, in view of selecting among these. Little assumption is made on the collection of estimators. The procedure offers the possibility to perform model selection and also to select among estimators associated to different model selection strategies. Besides, it provides an alternative to the $T$-estimators as studied recently in Birgé (2006). For illustration, we consider the problems of estimation and (complete) variable selection in various regression settings.

math.ST

A Bernstein-type inequality for suprema of random processes with an application to statistics

We use the generic chaining device proposed by Talagrand to establish exponential bounds on the deviation probability of some suprema of random processes. Then, given a random vector $ξ$ in $\R^{n}$ the components of which are independent and admit a suitable exponential moment, we deduce a deviation inequality for the squared Euclidean norm of the projection of $ξ$ onto a linear subspace of $\R^{n}$. Finally, we provide an application of such an inequality to statistics, performing model selection in the regression setting when the errors are possibly non-Gaussian and the collection of models possibly large.

math.ST