SearcharxivSearch

arXiv subjects

Lukas Steinberger

Publications and source records attributed to Lukas Steinberger.

13 recordsLinked to original sources

Implicit vs. explicit regularization for high-dimensional gradient descent

In this paper we investigate the generalization error of gradient descent (GD) applied to an $\ell_2$-regularized OLS objective function in the linear model. Based on our analysis we develop new methodology for computationally tractable and statistically efficient linear prediction in a high-dimensional and massive data scenario (large-$n$, large-$p$). Our results are based on the surprising observation that the generalization error of optimally tuned regularized gradient descent approaches that of an optimal benchmark procedure $monotonically$ in the iteration number $t$. On the other hand standard GD for OLS (without explicit regularization) can achieve the benchmark only in degenerate cases. This shows that (optimal) explicit regularization can be nearly statistically efficient (for large $t$) whereas implicit regularization by (optimal) early stopping can not. To complete our methodology, we provide a fully data driven and computationally tractable choice of the $\ell_2$ regularization parameter $λ$ that is computationally cheaper than cross-validation. On this way, we follow and extend ideas of Dicker (2014) to the non-gaussian case, which requires new results on high-dimensional sample covariance matrices that might be of independent interest.

math.ST

Towards multi-purpose locally differentially-private synthetic data release via spline wavelet plug-in estimation

We develop plug-in estimators for locally differentially private semi-parametric estimation via spline wavelets. The approach leads to optimal rates of convergence for a large class of estimation problems that are characterized by (differentiable) functionals $Λ(f)$ of the true data generating density $f$. The crucial feature of the locally private data $Z_1,\dots, Z_n$ we generate is that it does not depend on the particular functional $Λ$ (or the unknown density $f$) the analyst wants to estimate. Hence, the synthetic data can be generated and stored a priori and can subsequently be used by any number of analysts to estimate many vastly different functionals of interest at the provably optimal rate. In principle, this removes a long standing practical limitation in statistics of differential privacy, namely, that optimal privacy mechanisms need to be tailored towards the specific estimation problem at hand.

math.ST

Uncertainty quantification via cross-validation and its variants under algorithmic stability

Recently, there has been substantial interest in statistical guarantees for cross-validation (CV) methods of uncertainty quantification in statistical learning (cf. Barber et al. 2021a, Liang and Barber 2024, Steinberger and Leeb 2023). These guarantees should hold under minimal assumptions on the data generating process and conditional on the training data, because numerous predictions are usually computed based on one and the same training sample. We push this objective to the limit: We prove asymptotic conditional conservativeness of CV, that is, the probability of the actual coverage probability, conditional on the training data, undershooting its nominal level vanishes asymptotically, under minimal assumptions. In particular, we impose a stability condition, require that the prediction error is stochastically bounded, and show that neither condition can be dropped in general. By way of an asymptotic equivalence result, we also show that the closely related CV+ method of Barber et al. (2021a) provides exactly the same conditional statistical guarantees as CV in large samples, thereby extending the range of applicability of CV+ to the high-dimensional regime. We conclude that, in view of its marginal coverage guarantee, CV+ does indeed improve over simple CV. For our proofs we introduce a new concept called Lévy gauge, which can be of independent interest.

math.ST

Efficient Estimation of a Gaussian Mean with Local Differential Privacy

In this paper we study the problem of estimating the unknown mean $θ$ of a unit variance Gaussian distribution in a locally differentially private (LDP) way. In the high-privacy regime ($ε\le 1$), we identify an optimal privacy mechanism that minimizes the variance of the estimator asymptotically. Our main technical contribution is the maximization of the Fisher-Information of the sanitized data with respect to the local privacy mechanism $Q$. We find that the exact solution $Q_{θ,ε}$ of this maximization is the sign mechanism that applies randomized response to the sign of $X_i-θ$, where $X_1,\dots, X_n$ are the confidential iid original samples. However, since this optimal local mechanism depends on the unknown mean $θ$, we employ a two-stage LDP parameter estimation procedure which requires splitting agents into two groups. The first $n_1$ observations are used to consistently but not necessarily efficiently estimate the parameter $θ$ by $\tildeθ_{n_1}$. Then this estimate is updated by applying the sign mechanism with $\tildeθ_{n_1}$ instead of $θ$ to the remaining $n-n_1$ observations, to obtain an LDP and efficient estimator of the unknown mean.

math.ST

Efficiency in local differential privacy

We develop a theory of asymptotic efficiency in regular parametric models when data confidentiality is ensured by local differential privacy (LDP). Even though efficient parameter estimation is a classical and well-studied problem in mathematical statistics, it leads to several non-trivial obstacles that need to be tackled when dealing with the LDP case. Starting from a standard parametric model $\mathcal P=(P_θ)_{θ\inΘ}$, $Θ\subseteq\mathbb R^p$, for the iid unobserved sensitive data $X_1,\dots, X_n$, we establish local asymptotic mixed normality (along subsequences) of the model $$Q^{(n)}\mathcal P=(Q^{(n)}P_θ^n)_{θ\inΘ}$$ generating the sanitized observations $Z_1,\dots, Z_n$, where $Q^{(n)}$ is an arbitrary sequence of sequentially interactive privacy mechanisms. This result readily implies convolution and local asymptotic minimax theorems. In case $p=1$, the optimal asymptotic variance is found to be the inverse of the supremal Fisher-Information $\sup_{Q\in\mathcal Q_α} I_θ(Q\mathcal P)\in\mathbb R$, where the supremum runs over all $α$-differentially private (marginal) Markov kernels. We present an algorithm for finding a (nearly) optimal privacy mechanism $\hat{Q}$ and an estimator $\hatθ_n(Z_1,\dots, Z_n)$ based on the corresponding sanitized data that achieves this asymptotically optimal variance.

math.ST

Interactive versus non-interactive locally differentially private estimation: Two elbows for the quadratic functional

Local differential privacy has recently received increasing attention from the statistics community as a valuable tool to protect the privacy of individual data owners without the need of a trusted third party. Similar to the classical notion of randomized response, the idea is that data owners randomize their true information locally and only release the perturbed data. Many different protocols for such local perturbation procedures can be designed. In most estimation problems studied in the literature so far, however, no significant difference in terms of minimax risk between purely non-interactive protocols and protocols that allow for some amount of interaction between individual data providers could be observed. In this paper we show that for estimating the integrated square of a density, sequentially interactive procedures improve substantially over the best possible non-interactive procedure in terms of minimax rate of estimation. In particular, in the non-interactive scenario we identify an elbow in the minimax rate at $s=\frac34$, whereas in the sequentially interactive scenario the elbow is at $s=\frac12$. This is markedly different from both, the case of direct observations, where the elbow is well known to be at $s=\frac14$, as well as from the case where Laplace noise is added to the original data, where an elbow at $s= \frac94$ is obtained. We also provide adaptive estimators that achieve the optimal rate up to log-factors, we draw connections to non-parametric goodness-of-fit testing and estimation of more general integral functionals and conduct a series of numerical experiments. The fact that a particular locally differentially private, but interactive, mechanism improves over the simple non-interactive one is also of great importance for practical implementations of local differential privacy.

math.ST

Conditional predictive inference for stable algorithms

We investigate generically applicable and intuitively appealing prediction intervals based on $k$-fold cross validation. We focus on the conditional coverage probability of the proposed intervals, given the observations in the training sample (hence, training conditional validity), and show that it is close to the nominal level, in an appropriate sense, provided that the underlying algorithm used for computing point predictions is sufficiently stable when feature-response pairs are omitted. Our results are based on a finite sample analysis of the empirical distribution function of $k$-fold cross validation residuals and hold in non-parametric settings with only minimal assumptions on the error distribution. To illustrate our results, we also apply them to high-dimensional linear predictors, where we obtain uniform asymptotic training conditional validity as both sample size and dimension tend to infinity at the same rate and consistent parameter estimation typically fails. These results show that despite the serious problems of resampling procedures for inference on the unknown parameters (cf. Bickel and Freedman, 1983; El Karoui and Purdom, 2018; Mammen, 1996), cross validation methods can be successfully applied to obtain reliable predictive inference even in high dimensions and conditionally on the training data.

math.ST

Geometrizing rates of convergence under local differential privacy constraints

We study the problem of estimating a functional $θ(\mathbb P)$ of an unknown probability distribution $\mathbb P \in\mathcal P$ in which the original iid sample $X_1,\dots, X_n$ is kept private even from the statistician via an $α$-local differential privacy constraint. Let $ω_{TV}$ denote the modulus of continuity of the functional $θ$ over $\mathcal P$, with respect to total variation distance. For a large class of loss functions $l$ and a fixed privacy level $α$, we prove that the privatized minimax risk is equivalent to $l(ω_{TV}(n^{-1/2}))$ to within constants, under regularity conditions that are satisfied, in particular, if $θ$ is linear and $\mathcal P$ is convex. Our results complement the theory developed by Donoho and Liu (1991) with the nowadays highly relevant case of privatized data. Somewhat surprisingly, the difficulty of the estimation problem in the private case is characterized by $ω_{TV}$, whereas, it is characterized by the Hellinger modulus of continuity if the original data $X_1,\dots, X_n$ are available. We also find that for locally private estimation of linear functionals over a convex model a simple sample mean estimator, based on independently and binary privatized observations, always achieves the minimax rate. We further provide a general recipe for choosing the functional parameter in the optimal binary privatization mechanisms and illustrate the general theory in numerous examples. Our theory allows to quantify the price to be paid for local differential privacy in a large class of estimation problems. This price appears to be highly problem specific.

math.ST

Statistical inference with F-statistics when fitting simple models to high-dimensional data

We study linear subset regression in the context of the high-dimensional overall model $y = \vartheta+θ' z + ε$ with univariate response $y$ and a $d$-vector of random regressors $z$, independent of $ε$. Here, "high-dimensional" means that the number $d$ of available explanatory variables is much larger than the number $n$ of observations. We consider simple linear sub-models where $y$ is regressed on a set of $p$ regressors given by $x = M'z$, for some $d \times p$ matrix $M$ of full rank $p < n$. The corresponding simple model, i.e., $y=α+β' x + e$, can be justified by imposing appropriate restrictions on the unknown parameter $θ$ in the overall model; otherwise, this simple model can be grossly misspecified. In this paper, we establish asymptotic validity of the standard $F$-test on the surrogate parameter $β$, in an appropriate sense, even when the simple model is misspecified.

math.ST

Uniformly valid confidence intervals post-model-selection

We suggest general methods to construct asymptotically uniformly valid confidence intervals post-model-selection. The constructions are based on principles recently proposed by Berk et al. (2013). In particular the candidate models used can be misspecified, the target of inference is model-specific, and coverage is guaranteed for any data-driven model selection procedure. After developing a general theory we apply our methods to practically important situations where the candidate set of models, from which a working model is selected, consists of fixed design homoskedastic or heteroskedastic linear models, or of binary regression models with general link functions. In an extensive simulation study, we find that the proposed confidence intervals perform remarkably well, even when compared to existing methods that are tailored only for specific model selection procedures.

math.ST

The relative effects of dimensionality and multiplicity of hypotheses on the F-test in linear regression

Recently, several authors have re-examined the power of the classical F-test in linear regression in a `large-p, large-n' framework (cf. Zhong and Chen (2011), Wang and Cui (2013)). They highlight the loss of power as the number of regressors p increases relative to sample size n. These papers essentially focus only on the overall test of the null hypothesis that all p slope coefficients are equal to zero. Here, we consider the general case of testing q linear hypotheses on the (p+1)-dimensional regression parameter vector that includes p slope coefficients and an intercept parameter. In the case of Gaussian design, we describe the dependence of the local asymptotic power function on both the relative number of parameters p and the number of hypotheses q being tested, showing that the negative effect of dimensionality is less severe if the number of hypotheses is small. Using the recent work of Srivastava and Vershynin (2013) on high dimensional sample covariance matrices we are also able to substantially generalize previous results for non-Gaussian regressors.

math.ST

On conditional moments of high-dimensional random vectors given lower-dimensional projections

One of the most widely used properties of the multivariate Gaussian distribution, besides its tail behavior, is the fact that conditional means are linear and that conditional variances are constant. We here show that this property is also shared, in an approximate sense, by a large class of non-Gaussian distributions. We allow for several conditioning variables and we provide explicit non-asymptotic results, whereby we extend earlier findings of Hall and Li (1993) and Leeb (2013).

math.ST

Leave-one-out prediction intervals in linear regression models with many variables

We study prediction intervals based on leave-one-out residuals in a linear regression model where the number of explanatory variables can be large compared to sample size. We establish uniform asymptotic validity (conditional on the training sample) of the proposed interval under minimal assumptions on the unknown error distribution and the high dimensional design. Our intervals are generic in the sense that they are valid for a large class of linear predictors used to obtain a point forecast, such as robust M-estimators, James-Stein type estimators and penalized estimators like the LASSO. These results show that despite the serious problems of resampling procedures for inference on the unknown parameters, leave-one-out methods can be successfully applied to obtain reliable predictive inference even in high dimensions.

math.ST