SearcharxivSearch

arXiv subjects

Mayya Zhilova

Publications and source records attributed to Mayya Zhilova.

6 recordsLinked to original sources

New Edgeworth-type expansions with finite sample guarantees

We establish higher-order nonasymptotic expansions for a difference between probability distributions of sums of i.i.d. random vectors in a Euclidean space. The derived bounds are uniform over two classes of sets: the set of all Euclidean balls and the set of all half-spaces. These results allow to account for an impact of higher-order moments or cumulants of the considered distributions; the obtained error terms depend on a sample size and a dimension explicitly. The new inequalities outperform accuracy of the normal approximation in existing Berry-Esseen inequalities under very general conditions. Under some symmetry assumptions on the probability distribution of random summands, the obtained results are optimal in terms of the ratio between the dimension and the sample size. The new technique which we developed for establishing nonasymptotic higher-order expansions can be interesting by itself. Using the new higher-order inequalities, we study accuracy of the nonparametric bootstrap approximation and propose a bootstrap score test under possible model misspecification. The results of the paper also include explicit error bounds for general elliptic confidence regions for an expected value of the random summands, and optimality of the Gaussian anti-concentration inequality over the set of all Euclidean balls.

math.ST

Nonclassical Berry-Esseen inequalities and accuracy of the bootstrap

We study accuracy of bootstrap procedures for estimation of quantiles of a smooth function of a sum of independent sub-Gaussian random vectors. We establish higher-order approximation bounds with error terms depending on a sample size and a dimension explicitly. These results lead to improvements of accuracy of a weighted bootstrap procedure for general log-likelihood ratio statistics. The key element of our proofs of the bootstrap accuracy is a multivariate higher-order Berry-Esseen inequality. We consider a problem of approximation of distributions of two sums of zero mean independent random vectors, such that summands with the same indices have equal moments up to at least the second order. The derived approximation bound is uniform on the sets of all Euclidean balls. The presented approach extends classical Berry-Esseen type inequalities to higher-order approximation bounds. The theoretical results are illustrated with numerical experiments.

math.ST

Estimation of Smooth Functionals in Normal Models: Bias Reduction and Asymptotic Efficiency

Let $X_1,\dots, X_n$ be i.i.d. random variables sampled from a normal distribution $N(μ,Σ)$ in ${\mathbb R}^d$ with unknown parameter $θ=(μ,Σ)\in Θ:={\mathbb R}^d\times {\mathcal C}_+^d,$ where ${\mathcal C}_+^d$ is the cone of positively definite covariance operators in ${\mathbb R}^d.$ Given a smooth functional $f:Θ\mapsto {\mathbb R}^1,$ the goal is to estimate $f(θ)$ based on $X_1,\dots, X_n.$ Let $$ Θ(a;d):={\mathbb R}^d\times \Bigl\{Σ\in {\mathcal C}_+^d: σ(Σ)\subset [1/a, a]\Bigr\}, a\geq 1, $$ where $σ(Σ)$ is the spectrum of covariance $Σ.$ Let $\hat θ:=(\hat μ, \hat Σ),$ where $\hat μ$ is the sample mean and $\hat Σ$ is the sample covariance, based on the observations $X_1,\dots, X_n.$ For an arbitrary functional $f\in C^s(Θ),$ $s=k+1+ρ, k\geq 0, ρ\in (0,1],$ we define a functional $f_k:Θ\mapsto {\mathbb R}$ such that \begin{align*} & \sup_{θ\in Θ(a;d)}\|f_k(\hat θ)-f(θ)\|_{L_2({\mathbb P}_θ)} \lesssim_{s, β} \|f\|_{C^{s}(Θ)} \biggr[\biggl(\frac{a}{\sqrt{n}} \bigvee a^{βs}\biggl(\sqrt{\frac{d}{n}}\biggr)^{s} \biggr)\wedge 1\biggr], \end{align*} where $β=1$ for $k=0$ and $β>s-1$ is arbitrary for $k\geq 1.$ This error rate is minimax optimal and similar bounds hold for more general loss functions. If $d=d_n\leq n^α$ for some $α\in (0,1)$ and $s\geq \frac{1}{1-α},$ the rate becomes $O(n^{-1/2}).$ Moreover, for $s>\frac{1}{1-α},$ the estimators $f_k(\hat θ)$ is shown to be asymptotically efficient. The crucial part of the construction of estimator $f_k(\hat θ)$ is a bias reduction method studied in the paper for more general statistical models than normal.

math.ST

Efficient Estimation of Smooth Functionals in Gaussian Shift Models

We study a problem of estimation of smooth functionals of parameter $θ$ of Gaussian shift model $$ X=θ+ξ,\ θ\in E, $$ where $E$ is a separable Banach space and $X$ is an observation of unknown vector $θ$ in Gaussian noise $ξ$ with zero mean and known covariance operator $Σ.$ In particular, we develop estimators $T(X)$ of $f(θ)$ for functionals $f:E\mapsto {\mathbb R}$ of Hölder smoothness $s>0$ such that $$ \sup_{\|θ\|\leq 1} {\mathbb E}_θ(T(X)-f(θ))^2 \lesssim \Bigl(\|Σ\| \vee ({\mathbb E}\|ξ\|^2)^s\Bigr)\wedge 1, $$ where $\|Σ\|$ is the operator norm of $Σ,$ and show that this mean squared error rate is minimax optimal at least in the case of standard Gaussian shift model ($E={\mathbb R}^d$ equipped with the canonical Euclidean norm, $ξ=σZ,$ $Z\sim {\mathcal N}(0;I_d)$). Moreover, we determine a sharp threshold on the smoothness $s$ of functional $f$ such that, for all $s$ above the threshold, $f(θ)$ can be estimated efficiently with a mean squared error rate of the order $\|Σ\|$ in a "small noise" setting (that is, when ${\mathbb E}\|ξ\|^2$ is small). The construction of efficient estimators is crucially based on a "bootstrap chain" method of bias reduction. The results could be applied to a variety of special high-dimensional and infinite-dimensional Gaussian models (for vector, matrix and functional data).

math.ST

Bootstrap confidence sets under model misspecification

A multiplier bootstrap procedure for construction of likelihood-based confidence sets is considered for finite samples and a possible model misspecification. Theoretical results justify the bootstrap validity for a small or moderate sample size and allow to control the impact of the parameter dimension $p$: the bootstrap approximation works if $p^3/n$ is small. The main result about bootstrap validity continues to apply even if the underlying parametric model is misspecified under the so-called small modelling bias condition. In the case when the true model deviates significantly from the considered parametric family, the bootstrap procedure is still applicable but it becomes a bit conservative: the size of the constructed confidence sets is increased by the modelling bias. We illustrate the results with numerical examples for misspecified linear and logistic regressions.

math.ST

Simultaneous likelihood-based bootstrap confidence sets for a large number of models

The paper studies a problem of constructing simultaneous likelihood-based confidence sets. We consider a simultaneous multiplier bootstrap procedure for estimating the quantiles of the joint distribution of the likelihood ratio statistics, and for adjusting the confidence level for multiplicity. Theoretical results state the bootstrap validity in the following setting: the sample size \(n\) is fixed, the maximal parameter dimension \(p_{\textrm{max}}\) and the number of considered parametric models \(K\) are s.t. \((\log K)^{12}p_{\max}^{3}/n\) is small. We also consider the situation when the parametric models are misspecified. If the models' misspecification is significant, then the bootstrap critical values exceed the true ones and the simultaneous bootstrap confidence set becomes conservative. Numerical experiments for local constant and local quadratic regressions illustrate the theoretical results.

math.ST