Searcharxiv⌕ Search

arXiv subjects

Vladimir Koltchinskii

Publications and source records attributed to Vladimir Koltchinskii.

At least 19 recordsLinked to original sources

Estimation of trace functionals and spectral measures of covariance operators in Gaussian models

Let $f:{\mathbb R}_+\mapsto {\mathbb R}$ be a smooth function with $f(0)=0.$ A problem of estimation of a functional $τ_f(Σ):= {\rm tr}(f(Σ))$ of unknown covariance operator $Σ$ in a separable Hilbert space ${\mathbb H}$ based on i.i.d. mean zero Gaussian observations $X_1,\dots, X_n$ with values in ${\mathbb H}$ and covariance operator $Σ$ is studied. Let $\hat Σ_n$ be the sample covariance operator based on observations $X_1,\dots, X_n.$ Estimators \begin{align*} T_{f,m}(X_1,\dots, X_n):= \sum_{j=1}^m C_j τ_f(\hat Σ_{n_j}) \end{align*} based on linear aggregation of several plug-in estimators $τ_f(\hat Σ_{n_j}),$ where the sample sizes $n/c\leq n_1<\dots<n_m\leq n$ and coefficients $C_1,\dots, C_n$ are chosen to reduce the bias, are considered. The complexity of the problem is characterized by the effective rank ${\bf r}(Σ):= \frac{{\rm tr}(Σ)}{\|Σ\|}$ of covariance operator $Σ.$ It is shown that, if $f\in C^{m+1}({\mathbb R}_+)$ for some $m\geq 2,$ $\|f''\|_{L_{\infty}}\lesssim 1,$ $\|f^{(m+1)}\|_{L_{\infty}}\lesssim 1,$ $\|Σ\|\lesssim 1$ and ${\bf r}(Σ)\lesssim n,$ then \begin{align*} & \|\hat T_{f,m}(X_1,\dots, X_n)-τ_f(Σ)\|_{L_2} \lesssim_m \frac{\|Σf'(Σ)\|_2}{\sqrt{n}} + \frac{{\bf r}(Σ)}{n}+ {\bf r}(Σ)\Bigl(\sqrt{\frac{{\bf r}(Σ)}{n}}\Bigr)^{m+1}. \end{align*} Similar bounds have been proved for the $L_{p}$-errors and some other Orlicz norm errors of estimator $\hat T_{f,m}(X_1,\dots, X_n).$ The optimality of these error rates, other estimators for which asymptotic efficiency is achieved and uniform bounds over classes of smooth test functions $f$ are also discussed.

math.ST↗

Functional estimation in high-dimensional and infinite-dimensional models

Let ${\mathcal P}$ be a family of probability measures on a measurable space $(S,{\mathcal A}).$ Given a Banach space $E,$ a functional $f:E\mapsto {\mathbb R}$ and a mapping $θ: {\mathcal P}\mapsto E,$ our goal is to estimate $f(θ(P))$ based on i.i.d. observations $X_1,\dots, X_n\sim P, P\in {\mathcal P}.$ In particular, if ${\mathcal P}=\{P_θ: θ\in Θ\}$ is an identifiable statistical model with parameter set $Θ\subset E,$ one can consider the mapping $θ(P)=θ$ for $P\in {\mathcal P}, P=P_θ,$ resulting in a problem of estimation of $f(θ)$ based on i.i.d. observations $X_1,\dots, X_n\sim P_θ, θ\in Θ.$ Given a smooth functional $f$ and estimators $\hat θ_n(X_1,\dots, X_n), n\geq 1$ of $θ(P),$ we use these estimators, the sample split and the Taylor expansion of $f(θ(P))$ of a proper order to construct estimators $T_f(X_1,\dots, X_n)$ of $f(θ(P)).$ For these estimators and for a functional $f$ of smoothness $s\geq 1,$ we prove upper bounds on the $L_p$-errors of estimator $T_f(X_1,\dots, X_n)$ under certain moment assumptions on the base estimators $\hat θ_n.$ We study the performance of estimators $T_f(X_1,\dots, X_n)$ in several concrete problems, showing their minimax optimality and asymptotic efficiency. In particular, this includes functional estimation in high-dimensional models with many low dimensional components, functional estimation in high-dimensional exponential families and estimation of functionals of covariance operators in infinite-dimensional subgaussian models.

math.ST↗

Estimation of smooth functionals of covariance operators: jackknife bias reduction and bounds in terms of effective rank

Let $E$ be a separable Banach space and let $X, X_1,\dots, X_n, \dots$ be i.i.d. Gaussian random variables taking values in $E$ with mean zero and unknown covariance operator $Σ: E^{\ast}\mapsto E.$ The complexity of estimation of $Σ$ based on observations $X_1,\dots, X_n$ is naturally characterized by the so called effective rank of $Σ:$ ${\bf r}(Σ):= \frac{{\mathbb E}_Σ\|X\|^2}{\|Σ\|},$ where $\|Σ\|$ is the operator norm of $Σ.$ Given a smooth real valued functional $f$ defined on the space $L(E^{\ast},E)$ of symmetric linear operators from $E^{\ast}$ into $E$ (equipped with the operator norm), our goal is to study the problem of estimation of $f(Σ)$ based on $X_1,\dots, X_n.$ The estimators of $f(Σ)$ based on jackknife type bias reduction are considered and the dependence of their Orlicz norm error rates on effective rank ${\bf r}(Σ),$ the sample size $n$ and the degree of Hölder smoothness $s$ of functional $f$ are studied. In particular, it is shown that, if ${\bf r}(Σ)\lesssim n^α$ for some $α\in (0,1)$ and $s\geq \frac{1}{1-α},$ then the classical $\sqrt{n}$-rate is attainable and, if $s> \frac{1}{1-α},$ then asymptotic normality and asymptotic efficiency of the resulting estimators hold. Previously, the results of this type (for different estimators) were obtained only in the case of finite dimensional Euclidean space $E={\mathbb R}^d$ and for covariance operators $Σ$ whose spectrum is bounded away from zero (in which case, ${\bf r}(Σ)\asymp d$).

math.ST↗

Estimation of smooth functionals in high-dimensional models: bootstrap chains and Gaussian approximation

Let $X^{(n)}$ be an observation sampled from a distribution $P_θ^{(n)}$ with an unknown parameter $θ,$ $θ$ being a vector in a Banach space $E$ (most often, a high-dimensional space of dimension $d$). We study the problem of estimation of $f(θ)$ for a functional $f:E\mapsto {\mathbb R}$ of some smoothness $s>0$ based on an observation $X^{(n)}\sim P_θ^{(n)}.$ Assuming that there exists an estimator $\hat θ_n=\hat θ_n(X^{(n)})$ of parameter $θ$ such that $\sqrt{n}(\hat θ_n-θ)$ is sufficiently close in distribution to a mean zero Gaussian random vector in $E,$ we construct a functional $g:E\mapsto {\mathbb R}$ such that $g(\hat θ_n)$ is an asymptotically normal estimator of $f(θ)$ with $\sqrt{n}$ rate provided that $s>\frac{1}{1-α}$ and $d\leq n^α$ for some $α\in (0,1).$ We also derive general upper bounds on Orlicz norm error rates for estimator $g(\hat θ)$ depending on smoothness $s,$ dimension $d,$ sample size $n$ and the accuracy of normal approximation of $\sqrt{n}(\hat θ_n-θ).$ In particular, this approach yields asymptotically efficient estimators in some high-dimensional exponential models.

math.ST↗

Functional estimation in log-concave location families

Let $\{P_θ:θ\in {\mathbb R}^d\}$ be a log-concave location family with $P_θ(dx)=e^{-V(x-θ)}dx,$ where $V:{\mathbb R}^d\mapsto {\mathbb R}$ is a known convex function and let $X_1,\dots, X_n$ be i.i.d. r.v. sampled from distribution $P_θ$ with an unknown location parameter $θ.$ The goal is to estimate the value $f(θ)$ of a smooth functional $f:{\mathbb R}^d\mapsto {\mathbb R}$ based on observations $X_1,\dots, X_n.$ In the case when $V$ is sufficiently smooth and $f$ is a functional from a ball in a Hölder space $C^s,$ we develop estimators of $f(θ)$ with minimax optimal error rates measured by the $L_2({\mathbb P}_θ)$-distance as well as by more general Orlicz norm distances. Moreover, we show that if $d\leq n^α$ and $s>\frac{1}{1-α},$ then the resulting estimators are asymptotically efficient in Hájek-LeCam sense with the convergence rate $\sqrt{n}.$ This generalizes earlier results on estimation of smooth functionals in Gaussian shift models. The estimators have the form $f_k(\hat θ),$ where $\hat θ$ is the maximum likelihood estimator and $f_k: {\mathbb R}^d\mapsto {\mathbb R}$ (with $k$ depending on $s$) are functionals defined in terms of $f$ and designed to provide a higher order bias reduction in functional estimation problem. The method of bias reduction is based on iterative parametric bootstrap and it has been successfully used before in the case of Gaussian models.

math.ST↗

Estimation of Smooth Functionals in Normal Models: Bias Reduction and Asymptotic Efficiency

Let $X_1,\dots, X_n$ be i.i.d. random variables sampled from a normal distribution $N(μ,Σ)$ in ${\mathbb R}^d$ with unknown parameter $θ=(μ,Σ)\in Θ:={\mathbb R}^d\times {\mathcal C}_+^d,$ where ${\mathcal C}_+^d$ is the cone of positively definite covariance operators in ${\mathbb R}^d.$ Given a smooth functional $f:Θ\mapsto {\mathbb R}^1,$ the goal is to estimate $f(θ)$ based on $X_1,\dots, X_n.$ Let $$ Θ(a;d):={\mathbb R}^d\times \Bigl\{Σ\in {\mathcal C}_+^d: σ(Σ)\subset [1/a, a]\Bigr\}, a\geq 1, $$ where $σ(Σ)$ is the spectrum of covariance $Σ.$ Let $\hat θ:=(\hat μ, \hat Σ),$ where $\hat μ$ is the sample mean and $\hat Σ$ is the sample covariance, based on the observations $X_1,\dots, X_n.$ For an arbitrary functional $f\in C^s(Θ),$ $s=k+1+ρ, k\geq 0, ρ\in (0,1],$ we define a functional $f_k:Θ\mapsto {\mathbb R}$ such that \begin{align*} & \sup_{θ\in Θ(a;d)}\|f_k(\hat θ)-f(θ)\|_{L_2({\mathbb P}_θ)} \lesssim_{s, β} \|f\|_{C^{s}(Θ)} \biggr[\biggl(\frac{a}{\sqrt{n}} \bigvee a^{βs}\biggl(\sqrt{\frac{d}{n}}\biggr)^{s} \biggr)\wedge 1\biggr], \end{align*} where $β=1$ for $k=0$ and $β>s-1$ is arbitrary for $k\geq 1.$ This error rate is minimax optimal and similar bounds hold for more general loss functions. If $d=d_n\leq n^α$ for some $α\in (0,1)$ and $s\geq \frac{1}{1-α},$ the rate becomes $O(n^{-1/2}).$ Moreover, for $s>\frac{1}{1-α},$ the estimators $f_k(\hat θ)$ is shown to be asymptotically efficient. The crucial part of the construction of estimator $f_k(\hat θ)$ is a bias reduction method studied in the paper for more general statistical models than normal.

math.ST↗

Efficient Estimation of Smooth Functionals in Gaussian Shift Models

We study a problem of estimation of smooth functionals of parameter $θ$ of Gaussian shift model $$ X=θ+ξ,\ θ\in E, $$ where $E$ is a separable Banach space and $X$ is an observation of unknown vector $θ$ in Gaussian noise $ξ$ with zero mean and known covariance operator $Σ.$ In particular, we develop estimators $T(X)$ of $f(θ)$ for functionals $f:E\mapsto {\mathbb R}$ of Hölder smoothness $s>0$ such that $$ \sup_{\|θ\|\leq 1} {\mathbb E}_θ(T(X)-f(θ))^2 \lesssim \Bigl(\|Σ\| \vee ({\mathbb E}\|ξ\|^2)^s\Bigr)\wedge 1, $$ where $\|Σ\|$ is the operator norm of $Σ,$ and show that this mean squared error rate is minimax optimal at least in the case of standard Gaussian shift model ($E={\mathbb R}^d$ equipped with the canonical Euclidean norm, $ξ=σZ,$ $Z\sim {\mathcal N}(0;I_d)$). Moreover, we determine a sharp threshold on the smoothness $s$ of functional $f$ such that, for all $s$ above the threshold, $f(θ)$ can be estimated efficiently with a mean squared error rate of the order $\|Σ\|$ in a "small noise" setting (that is, when ${\mathbb E}\|ξ\|^2$ is small). The construction of efficient estimators is crucially based on a "bootstrap chain" method of bias reduction. The results could be applied to a variety of special high-dimensional and infinite-dimensional Gaussian models (for vector, matrix and functional data).

math.ST↗

Asymptotically Efficient Estimation of Smooth Functionals of Covariance Operators

Let $X$ be a centered Gaussian random variable in a separable Hilbert space ${\mathbb H}$ with covariance operator $Σ.$ We study a problem of estimation of a smooth functional of $Σ$ based on a sample $X_1,\dots ,X_n$ of $n$ independent observations of $X.$ More specifically, we are interested in functionals of the form $\langle f(Σ), B\rangle,$ where $f:{\mathbb R}\mapsto {\mathbb R}$ is a smooth function and $B$ is a nuclear operator in ${\mathbb H}.$ We prove concentration and normal approximation bounds for plug-in estimator $\langle f(\hat Σ),B\rangle,$ $\hat Σ:=n^{-1}\sum_{j=1}^n X_j\otimes X_j$ being the sample covariance based on $X_1,\dots, X_n.$ These bounds show that $\langle f(\hat Σ),B\rangle$ is an asymptotically normal estimator of its expectation ${\mathbb E}_Σ \langle f(\hat Σ),B\rangle$ (rather than of parameter of interest $\langle f(Σ),B\rangle$) with a parametric convergence rate $O(n^{-1/2})$ provided that the effective rank ${\bf r}(Σ):= \frac{{\bf tr}(Σ)}{\|Σ\|}$ (${\rm tr}(Σ)$ being the trace and $\|Σ\|$ being the operator norm of $Σ$) satisfies the assumption ${\bf r}(Σ)=o(n).$ At the same time, we show that the bias of this estimator is typically as large as $\frac{{\bf r}(Σ)}{n}$ (which is larger than $n^{-1/2}$ if ${\bf r}(Σ)\geq n^{1/2}$). In the case when ${\mathbb H}$ is finite-dimensional space of dimension $d=o(n),$ we develop a method of bias reduction and construct an estimator $\langle h(\hat Σ),B\rangle$ of $\langle f(Σ),B\rangle$ that is asymptotically normal with convergence rate $O(n^{-1/2}).$ Moreover, we study asymptotic properties of the risk of this estimator and prove minimax lower bounds for arbitrary estimators showing the asymptotic efficiency of $\langle h(\hat Σ),B\rangle$ in a semi-parametric sense.

math.ST↗

Efficient Estimation of Linear Functionals of Principal Components

We study principal component analysis (PCA) for mean zero i.i.d. Gaussian observations $X_1,\dots, X_n$ in a separable Hilbert space $\mathbb{H}$ with unknown covariance operator $Σ.$ The complexity of the problem is characterized by its effective rank ${\bf r}(Σ):= \frac{{\rm tr}(Σ)}{\|Σ\|},$ where ${\rm tr}(Σ)$ denotes the trace of $Σ$ and $\|Σ\|$ denotes its operator norm. We develop a method of bias reduction in the problem of estimation of linear functionals of eigenvectors of $Σ.$ Under the assumption that ${\bf r}(Σ)=o(n),$ we establish the asymptotic normality and asymptotic properties of the risk of the resulting estimators and prove matching minimax lower bounds, showing their semi-parametric optimality.

math.ST↗

Optimal Estimation of Low Rank Density Matrices

The density matrices are positively semi-definite Hermitian matrices of unit trace that describe the state of a quantum system. The goal of the paper is to develop minimax lower bounds on error rates of estimation of low rank density matrices in trace regression models used in quantum state tomography (in particular, in the case of Pauli measurements) with explicit dependence of the bounds on the rank and other complexity parameters. Such bounds are established for several statistically relevant distances, including quantum versions of Kullback-Leibler divergence (relative entropy distance) and of Hellinger distance (so called Bures distance), and Schatten $p$-norm distances. Sharp upper bounds and oracle inequalities for least squares estimator with von Neumann entropy penalization are obtained showing that minimax lower bounds are attained (up to logarithmic factors) for these distances.

stat.ML↗

Estimation of low rank density matrices: bounds in Schatten norms and other distances

Let ${\mathcal S}_m$ be the set of all $m\times m$ density matrices (Hermitian positively semi-definite matrices of unit trace). Consider a problem of estimation of an unknown density matrix $ρ\in {\mathcal S}_m$ based on outcomes of $n$ measurements of observables $X_1,\dots, X_n\in {\mathbb H}_m$ (${\mathbb H}_m$ being the space of $m\times m$ Hermitian matrices) for a quantum system identically prepared $n$ times in state $ρ.$ Outcomes $Y_1,\dots, Y_n$ of such measurements could be described by a trace regression model in which ${\mathbb E}_ρ(Y_j|X_j)={\rm tr}(ρX_j), j=1,\dots, n.$ The design variables $X_1,\dots, X_n$ are often sampled at random from the uniform distribution in an orthonormal basis $\{E_1,\dots, E_{m^2}\}$ of ${\mathbb H}_m$ (such as Pauli basis). The goal is to estimate the unknown density matrix $ρ$ based on the data $(X_1,Y_1), \dots, (X_n,Y_n).$ Let $$ \hat Z:=\frac{m^2}{n}\sum_{j=1}^n Y_j X_j $$ and let $\check ρ$ be the projection of $\hat Z$ onto the convex set ${\mathcal S}_m$ of density matrices. It is shown that for estimator $\check ρ$ the minimax lower bounds in classes of low rank density matrices (established earlier) are attained up logarithmic factors for all Schatten $p$-norm distances, $p\in [1,\infty]$ and for Bures version of quantum Hellinger distance. Moreover, for a slightly modified version of estimator $\check ρ$ the same property holds also for quantum relative entropy (Kullback-Leibler) distance between density matrices.

stat.ML↗

Nuclear norm penalization and optimal rates for noisy low rank matrix completion

This paper deals with the trace regression model where $n$ entries or linear combinations of entries of an unknown $m_1\times m_2$ matrix $A_0$ corrupted by noise are observed. We propose a new nuclear norm penalized estimator of $A_0$ and establish a general sharp oracle inequality for this estimator for arbitrary values of $n,m_1,m_2$ under the condition of isometry in expectation. Then this method is applied to the matrix completion problem. In this case, the estimator admits a simple explicit form and we prove that it satisfies oracle inequalities with faster rates of convergence than in the previous works. They are valid, in particular, in the high-dimensional setting $m_1m_2\gg n$. We show that the obtained rates are optimal up to logarithmic factors in a minimax sense and also derive, for any fixed matrix $A_0$, a non-minimax lower bound on the rate of convergence of our estimator, which coincides with the upper bound up to a constant factor. Finally, we show that our procedure provides an exact recovery of the rank of $A_0$ with probability close to 1. We also discuss the statistical learning setting where there is no underlying model determined by $A_0$ and the aim is to find the best trace regression model approximating the data.

math.ST↗

New asymptotic results in principal component analysis

Let $X$ be a mean zero Gaussian random vector in a separable Hilbert space ${\mathbb H}$ with covariance operator $Σ:={\mathbb E}(X\otimes X).$ Let $Σ=\sum_{r\geq 1}μ_r P_r$ be the spectral decomposition of $Σ$ with distinct eigenvalues $μ_1>μ_2> \dots$ and the corresponding spectral projectors $P_1, P_2, \dots.$ Given a sample $X_1,\dots, X_n$ of size $n$ of i.i.d. copies of $X,$ the sample covariance operator is defined as $\hat Σ_n := n^{-1}\sum_{j=1}^n X_j\otimes X_j.$ The main goal of principal component analysis is to estimate spectral projectors $P_1, P_2, \dots$ by their empirical counterparts $\hat P_1, \hat P_2, \dots$ properly defined in terms of spectral decomposition of the sample covariance operator $\hat Σ_n.$ The aim of this paper is to study asymptotic distributions of important statistics related to this problem, in particular, of statistic $\|\hat P_r-P_r\|_2^2,$ where $\|\cdot\|_2^2$ is the squared Hilbert--Schmidt norm. This is done in a "high-complexity" asymptotic framework in which the so called effective rank ${\bf r}(Σ):=\frac{{\rm tr}(Σ)}{\|Σ\|_{\infty}}$ (${\rm tr}(\cdot)$ being the trace and $\|\cdot\|_{\infty}$ being the operator norm) of the true covariance $Σ$ is becoming large simultaneously with the sample size $n,$ but ${\bf r}(Σ)=o(n)$ as $n\to\infty.$ In this setting, we prove that, in the case of one-dimensional spectral projector $P_r,$ the properly centered and normalized statistic $\|\hat P_r-P_r\|_2^2$ with {\it data-dependent} centering and normalization converges in distribution to a Cauchy type limit. The proofs of this and other related results rely on perturbation analysis and Gaussian concentration.

math.ST↗

Asymptotics and Concentration Bounds for Bilinear Forms of Spectral Projectors of Sample Covariance

Let $X,X_1,\dots, X_n$ be i.i.d. Gaussian random variables with zero mean and covariance operator $Σ={\mathbb E}(X\otimes X)$ taking values in a separable Hilbert space ${\mathbb H}.$ Let $$ {\bf r}(Σ):=\frac{{\rm tr}(Σ)}{\|Σ\|_{\infty}} $$ be the effective rank of $Σ,$ ${\rm tr}(Σ)$ being the trace of $Σ$ and $\|Σ\|_{\infty}$ being its operator norm. Let $$\hat Σ_n:=n^{-1}\sum_{j=1}^n (X_j\otimes X_j)$$ be the sample (empirical) covariance operator based on $(X_1,\dots, X_n).$ The paper deals with a problem of estimation of spectral projectors of the covariance operator $Σ$ by their empirical counterparts, the spectral projectors of $\hat Σ_n$ (empirical spectral projectors). The focus is on the problems where both the sample size $n$ and the effective rank ${\bf r}(Σ)$ are large. This framework includes and generalizes well known high-dimensional spiked covariance models. Given a spectral projector $P_r$ corresponding to an eigenvalue $μ_r$ of covariance operator $Σ$ and its empirical counterpart $\hat P_r,$ we derive sharp concentration bounds for bilinear forms of empirical spectral projector $\hat P_r$ in terms of sample size $n$ and effective dimension ${\bf r}(Σ).$ Building upon these concentration bounds, we prove the asymptotic normality of bilinear forms of random operators $\hat P_r -{\mathbb E}\hat P_r$ under the assumptions that $n\to \infty$ and ${\bf r}(Σ)=o(n).$ In a special case of eigenvalues of multiplicity one, these results are rephrased as concentration bounds and asymptotic normality for linear forms of empirical eigenvectors. Other results include bounds on the bias ${\mathbb E}\hat P_r-P_r$ and a method of bias reduction as well as a discussion of possible applications to statistical inference in high-dimensional principal component analysis.

math.ST↗

Perturbation of linear forms of singular vectors under Gaussian noise

Let $A\in\mathbb{R}^{m\times n}$ be a matrix of rank $r$ with singular value decomposition (SVD) $A=\sum_{k=1}^rσ_k (u_k\otimes v_k),$ where $\{σ_k, k=1,\ldots,r\}$ are singular values of $A$ (arranged in a non-increasing order) and $u_k\in {\mathbb R}^m, v_k\in {\mathbb R}^n, k=1,\ldots, r$ are the corresponding left and right orthonormal singular vectors. Let $\tilde{A}=A+X$ be a noisy observation of $A,$ where $X\in\mathbb{R}^{m\times n}$ is a random matrix with i.i.d. Gaussian entries, $X_{ij}\sim\mathcal{N}(0,τ^2),$ and consider its SVD $\tilde{A}=\sum_{k=1}^{m\wedge n}\tildeσ_k(\tilde{u}_k\otimes\tilde{v}_k)$ with singular values $\tildeσ_1\geq\ldots\geq\tildeσ_{m\wedge n}$ and singular vectors $\tilde{u}_k,\tilde{v}_k,k=1,\ldots, m\wedge n.$ The goal of this paper is to develop sharp concentration bounds for linear forms $\langle \tilde u_k,x\rangle, x\in {\mathbb R}^m$ and $\langle \tilde v_k,y\rangle, y\in {\mathbb R}^n$ of the perturbed (empirical) singular vectors in the case when the singular values of $A$ are distinct and, more generally, concentration bounds for bilinear forms of projection operators associated with SVD. In particular, the results imply upper bounds of the order $O\biggl(\sqrt{\frac{\log(m+n)}{m\vee n}}\biggr)$ (holding with a high probability) on $$\max_{1\leq i\leq m}\big|\big<\tilde{u}_k-\sqrt{1+b_k}u_k,e_i^m\big>\big|\ \ {\rm and} \ \ \max_{1\leq j\leq n}\big|\big<\tilde{v}_k-\sqrt{1+b_k}v_k,e_j^n\big>\big|,$$ where $b_k$ are properly chosen constants characterizing the bias of empirical singular vectors $\tilde u_k, \tilde v_k$ and $\{e_i^m,i=1,\ldots,m\}, \{e_j^n,j=1,\ldots,n\}$ are the canonical bases of $\mathbb{R}^m, {\mathbb R}^n,$ respectively.

math.PR↗

Normal approximation and concentration of spectral projectors of sample covariance

Let $X,X_1,\dots, X_n$ be i.i.d. Gaussian random variables in a separable Hilbert space ${\mathbb H}$ with zero mean and covariance operator $Σ={\mathbb E}(X\otimes X),$ and let $\hat Σ:=n^{-1}\sum_{j=1}^n (X_j\otimes X_j)$ be the sample (empirical) covariance operator based on $(X_1,\dots, X_n).$ Denote by $P_r$ the spectral projector of $Σ$ corresponding to its $r$-th eigenvalue $μ_r$ and by $\hat P_r$ the empirical counterpart of $P_r.$ The main goal of the paper is to obtain tight bounds on $$ \sup_{x\in {\mathbb R}} \left|{\mathbb P}\left\{\frac{\|\hat P_r-P_r\|_2^2-{\mathbb E}\|\hat P_r-P_r\|_2^2}{{\rm Var}^{1/2}(\|\hat P_r-P_r\|_2^2)}\leq x\right\}-Φ(x)\right|, $$ where $\|\cdot\|_2$ denotes the Hilbert--Schmidt norm and $Φ$ is the standard normal distribution function. Such accuracy of normal approximation of the distribution of squared Hilbert--Schmidt error is characterized in terms of so called effective rank of $Σ$ defined as ${\bf r}(Σ)=\frac{{\rm tr}(Σ)}{\|Σ\|_{\infty}},$ where ${\rm tr}(Σ)$ is the trace of $Σ$ and $\|Σ\|_{\infty}$ is its operator norm, as well as another parameter characterizing the size of ${\rm Var}(\|\hat P_r-P_r\|_2^2).$ Other results include non-asymptotic bounds and asymptotic representations for the mean squared Hilbert--Schmidt norm error ${\mathbb E}\|\hat P_r-P_r\|_2^2$ and the variance ${\rm Var}(\|\hat P_r-P_r\|_2^2),$ and concentration inequalities for $\|\hat P_r-P_r\|_2^2$ around its expectation.

math.ST↗

Estimation of Low-Rank Covariance Function

We consider the problem of estimating a low rank covariance function $K(t,u)$ of a Gaussian process $S(t), t\in [0,1]$ based on $n$ i.i.d. copies of $S$ observed in a white noise. We suggest a new estimation procedure adapting simultaneously to the low rank structure and the smoothness of the covariance function. The new procedure is based on nuclear norm penalization and exhibits superior performances as compared to the sample covariance function by a polynomial factor in the sample size $n$. Other results include a minimax lower bound for estimation of low-rank covariance functions showing that our procedure is optimal as well as a scheme to estimate the unknown noise variance of the Gaussian process.

math.ST↗

$L_1$-Penalization in Functional Linear Regression with Subgaussian Design

We study functional regression with random subgaussian design and real-valued response. The focus is on the problems in which the regression function can be well approximated by a functional linear model with the slope function being "sparse" in the sense that it can be represented as a sum of a small number of well separated "spikes". This can be viewed as an extension of now classical sparse estimation problems to the case of infinite dictionaries. We study an estimator of the regression function based on penalized empirical risk minimization with quadratic loss and the complexity penalty defined in terms of $L_1$-norm (a continuous version of LASSO). The main goal is to introduce several important parameters characterizing sparsity in this class of problems and to prove sharp oracle inequalities showing how the $L_2$-error of the continuous LASSO estimator depends on the underlying sparsity of the problem.

math.ST↗