SearcharxivSearch

arXiv subjects

Adrien Saumard

Publications and source records attributed to Adrien Saumard.

At least 19 recordsLinked to original sources

Pointwise convergence of purely random partition estimators: from random trees to prototype rules

We study pointwise convergence rates of purely random partition estimators in nonparametric regression, where the partition -- into hyper-rectangles by purely random trees, or into Voronoi cells by prototype rules -- is built independently of the responses. Our analysis rests on a single geometric criterion, shape regularity, relating the diameter of a cell to its volume, which is shown by Bettinger, Portier and Saumard (2026) to be necessary and sufficient, up to logarithmic factors, for achieving the minimax rate $n^{-1/(d+2)}$. We show that centered and uniform trees are not shape-regular -- their cells' aspect ratio grows exponentially with the number of splits with probability bounded away from zero -- explaining the super-logarithmic corrections in their error bounds, whereas Mondrian trees, whose splits adapt to the current cell geometry, are shape-regular in probability and attain the minimax rate. The same analysis applied to Voronoi partitions yields the first pointwise concentration bounds for Proto-NN, resolving an open problem of Gy\"orfi and Weiss (2021), and shows that OptiNet achieves the minimax rate with markedly better success probability -- even almost surely, for a suitable choice of parameters -- thanks to its $\eta$-net construction.

math.ST

Revisiting local regression: shape regularity, uniform rates, and the limits of random splits

Considering pointwise and sup-norm estimation, we analyze the non-asymptotic behavior of local averaging estimators for Lipschitz regression functions. Building on a general deviation bound for estimators based on a VC family of localizing sets, we introduce the notion of shape-regular local maps, where averaging is performed over sets with an almost isotropic geometry. Our main message is a characterization: shape regularity is both necessary and sufficient to attain optimal rates, up to logarithmic factors. Necessity is established non-asymptotically through an explicit anisotropic example, sharpening a phenomenon previously understood only heuristically in asymptotic theory. We then draw two consequences. First, the simple $k$-nearest neighbor rule is shape-regular by construction and attains the optimal rate, even on unbounded supports. Second, and perhaps surprisingly, the popular random-split condition for trees -- known to ensure consistency and vanishing cell diameters -- does not guarantee optimal rates: for blind tree constructions, the cell aspect ratio diverges exponentially with depth, so that shape regularity fails with positive probability. This identifies the absence of a geometric correction mechanism, rather than a slowly shrinking diameter, as the obstruction to optimality. Motivated by this gap, we propose a tree construction that enforces shape regularity through a simple constraint on admissible splits, and prove a uniform deviation inequality showing that it restores the optimal rate for Lipschitz functions.

math.ST

Concentration of the bootstrap empirical process, with applications to statistical inference

Considering a general framework of bootstrap with exchangeable weights, we show some concentration inequalities for the supremum of the bootstrap empirical process. On the one hand, we discuss the concentration of the bootstrap empirical process around its conditional expectation with respect to the original data, and on the other hand, the concentration of the latter quantity around its mean. For the concentration conditional on data, we build on Chatterjee's exchangeable pairs approach to concentration. To attain optimal concentration rates, we develop some refined arguments for the convergence of transposition walks on the symmetric group. The conditional expectation of the bootstrap empirical process is proved to be self-bounding, thus extending a well-known property for conditional Rademacher averages. To illustrate the interest of these concentration inequalities, we provide some new results pertaining to confidence regions for the estimation of a mean vector, as well as non-asymptotic bounds for the two-sample permutation test.

math.ST

On the pointwise and sup-norm errors for local regression estimators

In this paper, we analyze the behavior of various non-parametric local regression estimators, i.e. estimators that are based on local averaging, for estimating a Lipschitz regression function at a fixed point, or in sup-norm. We first prove some deviation bounds for local estimators that can be indexed by a VC class of sets in the covariates space. We then introduce the general concept of shape-regular local maps, corresponding to the situation where the local averaging is done on sets which, in some sense, have ``almost isotropic'' shapes. On the one hand, we prove that, in general, shape-regularity is necessary to achieve the minimax rates of convergence. On the other hand, we prove that it is sufficient to ensure the optimal rates, up to some logarithmic factors. Next, we prove some deviation bounds for specific estimators, that are based on data-dependent local maps, such as nearest neighbors, their recent prototype variants, as well as a new algorithm, which is a modified and generalized version of CART, and that is minimax rate optimal in sup-norm. In particular, the latter algorithm is based on a random tree construction that depends on both the covariates and the response data. For each of the estimators, we provide insights on the shape-regularity of their respective local maps. Finally, we conclude the paper by establishing some probability bounds for local estimators based on purely random trees, such as centered, uniform or Mondrian trees. Again, we discuss the relations between the rates of the estimators and the shape-regularity of their local maps.

math.ST

A theory of shape regularity for local regression maps

We introduce the concept of shape-regular regression maps as a framework to derive optimal rates of convergence for various non-parametric local regression estimators. Using Vapnik-Chervonenkis theory, we establish upper and lower bounds on the pointwise and the sup-norm estimation error, even when the localization procedure depends on the full data sample, and under mild conditions on the regression model. Our results demonstrate that the shape regularity of regression maps is not only sufficient but also necessary to achieve an optimal rate of convergence for Lipschitz regression functions. To illustrate the theory, we establish new concentration bounds for many popular local regression methods such as nearest neighbors algorithm, CART-like regression trees and several purely random trees including Mondrian trees.

math.ST

Covariance inequalities for convex and log-concave functions

Extending results of Harg{é} and Hu for the Gaussian measure, we prove inequalities for the covariance Cov$_μ(f, g)$ where $μ$ is a general product probability measure on $\mathbb{R}^d$ and $f,g: \mathbb{R}^d \to \mathbb{R}$ satisfy some convexity or log-concavity assumptions, with possibly some symmetries.

math.PR

Phase transitions for support recovery under local differential privacy

We address the problem of variable selection in a high-dimensional but sparse mean model, under the additional constraint that only privatised data are available for inference. The original data are vectors with independent entries having a symmetric, strongly log-concave distribution on $\mathbb{R}$. For this purpose, we adopt a recent generalisation of classical minimax theory to the framework of local $α-$differential privacy. We provide lower and upper bounds on the rate of convergence for the expected Hamming loss over classes of at most $s$-sparse vectors whose non-zero coordinates are separated from $0$ by a constant $a>0$. As corollaries, we derive necessary and sufficient conditions (up to log factors) for exact recovery and for almost full recovery. When we restrict our attention to non-interactive mechanisms that act independently on each coordinate our lower bound shows that, contrary to the non-private setting, both exact and almost full recovery are impossible whatever the value of $a$ in the high-dimensional regime such that $n α^2/ d^2\lesssim 1$. However, in the regime $nα^2/d^2\gg \log(d)$ we can exhibit a critical value $a^*$ (up to a logarithmic factor) such that exact and almost full recovery are possible for all $a\gg a^*$ and impossible for $a\leq a^*$. We show that these results can be improved when allowing for all non-interactive (that act globally on all coordinates) locally $α-$differentially private mechanisms in the sense that phase transitions occur at lower levels.

math.ST

Relaxing the Gaussian assumption in Shrinkage and SURE in high dimension

Shrinkage estimation is a fundamental tool of modern statistics, pioneered by Charles Stein upon his discovery of the famous paradox involving the multivariate Gaussian. A large portion of the subsequent literature only considers the efficiency of shrinkage, and that of an associated procedure known as Stein's Unbiased Risk Estimate, or SURE, in the Gaussian setting of that original work. We investigate what extensions to the domain of validity of shrinkage and SURE can be made away from the Gaussian through the use of tools developed in the probabilistic area now known as Stein's method. We show that shrinkage is efficient away from the Gaussian under very mild conditions on the distribution of the noise. SURE is also proved to be adaptive under similar assumptions, and in particular in a way that retains the classical asymptotics of Pinsker's theorem. Notably, shrinkage and SURE are shown to be efficient under mild distributional assumptions, and particularly for general isotropic log-concave measures.

math.ST

High-dimensional logistic entropy clustering

Minimization of the (regularized) entropy of classification probabilities is a versatile class of discriminative clustering methods. The classification probabilities are usually defined through the use of some classical losses from supervised classification and the point is to avoid modelisation of the full data distribution by just optimizing the law of the labels conditioned on the observations. We give the first theoretical study of such methods, by specializing to logistic classification probabilities. We prove that if the observations are generated from a two-component isotropic Gaussian mixture, then minimizing the entropy risk over a Euclidean ball indeed allows to identify the separation vector of the mixture. Furthermore, if this separation vector is sparse, then penalizing the empirical risk by a $\ell_{1}$-regularization term allows to infer the separation in a high-dimensional space and to recover its support, at standard rates of sparsity problems. Our approach is based on the local convexity of the logistic entropy risk, that occurs if the separation vector is large enough, with a condition on its norm that is independent from the space dimension. This local convexity property also guarantees fast rates in a classical, low-dimensional setting.

math.ST

K-bMOM: a robust Lloyd-type clustering algorithm based on bootstrap Median-of-Means

We propose a new clustering algorithm that is robust to the presence of outliers in the dataset. We perform Lloyd-type iterations with robust estimates of the centroids. More precisely, we build on the idea of median-of-means statistics to estimate the centroids, but allow for replacement while constructing the blocks. We call this methodology the bootstrap median-of-means (bMOM) and prove that if enough blocks are generated through the bootstrap sampling, then it has a better breakdown point for mean estimation than the classical median-of-means (MOM), where the blocks form a partition of the dataset. From a clustering perspective, bMOM enables to take many blocks of a desired size, thus avoiding possible disappearance of clusters in some blocks, a pitfall that can occur for the partition-based generation of blocks of the classical median-of-means. Experiments on simulated datasets show that the proposed approach, called K-bMOM, performs better than existing robust K-means based methods. Guidelines are provided for tuning the hyper-parameters K-bMOM in practice. It is also recommended to the practitionner to use such a robust approach to initialize their clustering algorithm. Finally, considering a simplified and theoretical version of our estimator, we prove its robustness to adversarial contamination by deriving robust rates of convergence for the K-means distorsion. To our knowledge, it is the first result of this kind for the K-means distorsion.

stat.ME

Bi-log-concavity: some properties and some remarks towards a multi-dimensional extension

Bi-log-concavity of probability measures is a univariate extension of the notion of log-concavity that has been recently proposed in a statistical literature. Among other things, it has the nice property from a modelisation perspective to admit some multimodal distributions, while preserving some nice features of log-concave measures. We compute the isoperimetric constant for a bi-log-concave measure, extending a property available for log-concave measures. This implies that bi-log-concave measures have exponentially decreasing tails. Then we show that the convolution of a bi-log-concave measure with a log-concave one is bi-log-concave. Consequently, infinitely differentiable, positive densities are dense in the set of bi-log-concave densities for $L_p-$norms, $p \in [1;+\infty]$. We also derive a necessary and sufficient condition for the convolution of two bi-log-concave measures to be bi-log-concave. We conclude this note by discussing ways of defining a multi-dimensional extension of the notion of bi-log-concavity. We propose an approach based on a variant of the isoperimetric problem, restricted to half-spaces.

math.PR

Local differential privacy: Elbow effect in optimal density estimation and adaptation over Besov ellipsoids

We address the problem of non-parametric density estimation under the additional constraint that only privatised data are allowed to be published and available for inference. For this purpose, we adopt a recent generalisation of classical minimax theory to the framework of local $α$-differential privacy and provide a lower bound on the rate of convergence over Besov spaces $B^s_{pq}$ under mean integrated $\mathbb L^r$-risk. This lower bound is deteriorated compared to the standard setup without privacy, and reveals a twofold elbow effect. In order to fulfil the privacy requirement, we suggest adding suitably scaled Laplace noise to empirical wavelet coefficients. Upper bounds within (at most) a logarithmic factor are derived under the assumption that $α$ stays bounded as $n$ increases: A linear but non-adaptive wavelet estimator is shown to attain the lower bound whenever $p \geq r$ but provides a slower rate of convergence otherwise. An adaptive non-linear wavelet estimator with appropriately chosen smoothing parameters and thresholding is shown to attain the lower bound within a logarithmic factor for all cases.

math.ST

Weighted Poincaré inequalities, concentration inequalities and tail bounds related to the Stein kernels in dimension one

We investigate links between the so-called Stein's density approach in dimension one and some functional and concentration inequalities. We show that measures having a finite first moment and a density with connected support satisfy a weighted Poincaré inequality with the weight being the Stein kernel, that indeed exists and is unique in this case. Furthermore, we prove weighted log-Sobolev and asymmetric Brascamp-Lieb type inequalities related to Stein kernels. We also show that existence of a uniformly bounded Stein kernel is sufficient to ensure a positive Cheeger isoperimetric constant. Then we derive new concentration inequalities. In particular, we prove generalized Mills' type inequalities when a Stein kernel is uniformly bounded and sub-gamma concentration for Lipschitz functions of a variable with a sub-linear Stein kernel. When some exponential moments are finite, a general concentration inequality is then expressed in terms of Legendre-Fenchel transform of the Laplace transform of the Stein kernel. Along the way, we prove a general lemma for bounding the Laplace transform of a random variable, that should be useful in many other contexts when deriving concentration inequalities. Finally, we provide density and tail formulas as well as tail bounds, generalizing previous results that where obtained in the context of Malliavin calculus.

math.PR

Finite sample improvement of Akaike's Information Criterion

We emphasize that it is possible to improve the principle of unbiased risk estimation for model selection by addressing excess risk deviations in the design of penalization procedures. Indeed, we propose a modification of Akaike's Information Criterion that avoids overfitting, even when the sample size is small. We call this correction an over-penalization procedure. As proof of concept, we show the nonasymptotic optimality of our histogram selection procedure in density estimation by establishing sharp oracle inequalities for the Kullback-Leibler divergence. One of the main features of our theoretical results is that they include the estimation of unbounded logdensities. To do so, we prove several analytical and probabilistic lemmas that are of independent interest. In an experimental study, we also demonstrate state-of-the-art performance of our over-penalization criterion for bin size selection, in particular outperforming AICc procedure.

math.ST

An extremal property of the normal distribution, with a discrete analog

We prove, using the Brascamp-Lieb inequality, that the Gaussian measure is the only strong log-concave measure having a strong log-concavity parameter equal to its covariance matrix. We also give a similar characterization of the Poisson measure in the discrete case, using "Chebyshev's other inequality". We briefly discuss how these results relate to Stein and Stein-Chen methods for Gaussian and Poisson approximation, and to the Bakry-Emery calculus.

math.PR

A concentration inequality for the excess risk in least-squares regression with random design and heteroscedastic noise

We prove a new and general concentration inequality for the excess risk in least-squares regression with random design and heteroscedastic noise. No specific structure is required on the model, except the existence of a suitable function that controls the local suprema of the empirical process. So far, only the case of linear contrast estimation was tackled in the literature with this level of generality on the model. We solve here the case of a quadratic contrast, by separating the behavior of a linearized empirical process and the empirical process driven by the squares of functions of models.

math.ST

On the isoperimetric constant, covariance inequalities and $L_p$-Poincaré inequalities in dimension one

Firstly, we derive in dimension one a new covariance inequality of $L_{1}-L_{\infty}$ type that characterizes the isoperimetric constant as the best constant achieving the inequality. Secondly, we generalize our result to $L_{p}-L_{q}$ bounds for the covariance. Consequently, we recover Cheeger's inequality without using the co-area formula. We also prove a generalized weighted Hardy type inequality that is needed to derive our covariance inequalities and that is of independent interest. Finally, we explore some consequences of our covariance inequalities for $L_{p}$-Poincaré inequalities and moment bounds. In particular, we obtain optimal constants in general $L_{p}$-Poincaré inequalities for measures with finite isoperimetric constant, thus generalizing in dimension one Cheeger's inequality, which is a $L_{p}$-Poincaré inequality for $p=2$, to any real $p\geq 1$.

math.PR

Efron's monotonicity property for measures on $\mathbb{R}^2$

First we prove some kernel representations for the covariance of two functions taken on the same random variable and deduce kernel representations for some functionals of a continuous one-dimensional measure. Then we apply these formulas to extend Efron's monotonicity property, given in Efron [1965] and valid for independent log-concave measures, to the case of general measures on $\mathbb{R}^2$. The new formulas are also used to derive some further quantitative estimates in Efron's monotonicity property.

math.ST