SearcharxivSearch

arXiv subjects

Vladimir Spokoiny

Publications and source records attributed to Vladimir Spokoiny.

At least 19 recordsLinked to original sources

High-Dimensional Change Point Detection via Graph Spanning Ratio

Inspired by graph-based methodologies, we introduce a novel graph-spanning algorithm designed to identify changes in both offline and online data across low to high dimensions. This versatile approach is applicable to Euclidean and graph-structured data with unknown distributions, while maintaining control over error probabilities. Theoretically, we demonstrate that the algorithm achieves high detection power when the magnitude of the change surpasses the lower bound of the minimax separation rate, which scales on the order of $\sqrt{nd}$. Our method outperforms other techniques in terms of accuracy for both Gaussian and non-Gaussian data. Notably, it maintains strong detection power even with small observation windows, making it particularly effective for online environments where timely and precise change detection is critical.

stat.ML

Generalized bootstrap in the Bures-Wasserstein space

This study proposes a bootstrap-based method for uncertainty quantification in two important statistical scenarios. First, we approximate the sampling distribution of empirical barycenters under the Bures--Wasserstein metric using a reweighted estimator. Our theoretical results guarantee the accuracy of this approximation and enable the construction of data-driven confidence sets. The methodology is validated through experiments on graph-structured data, including stochastic block models and brain connectomes. Additionally, we compare bootstrap-based confidence sets with the asymptotic confidence sets obtained in arXiv:1901.00226v2, evaluating both their statistical performance and computational complexity. Second, we investigate the generalized bootstrap framework for $M$-estimators without requiring a specific resampling scheme, thus covering both weighted and resampling methods under mild conditions. Both contributions rely on a novel Gaussian approximation result for $M$-estimators.

math.ST

Marginal minimization and sup-norm expansions in perturbed optimization

Let the objective unction \( f \) depends on the target variable \( x \) along with a nuisance variable \( s \): \( f(v) = f(x,s) \). The goal is to identify the marginal solution \( x^{*} = \arg\min_{x} \min_{s} f(x,s) \). This paper discusses three related problems. The plugin approach widely used e.g. in inverse problems suggests to use a preliminary guess (pilot) \( \hat{s} \) and apply the solution of the partial optimization \( \hat{x} = \arg\min_{x} f(x,\hat{s}) \). The main question to address within this approach is the required quality of the pilot ensuring the prescribed accuracy of \( \hat{x} \). The popular \emph{alternating optimization} approach suggests the following procedure: given a starting guess \( x_{0} \), for \( t \geq 1 \), define \( s_{t} = \arg\min_{s} f(x_{t-1},s) \), and then \( x_{t} = \arg\min_{x} f(x,s_{t}) \). The main question here is the set of conditions ensuring a convergence of \( x_{t} \) to \( x^{*} \). Finally, the paper discusses an interesting connection between marginal optimization and sup-norm estimation. The basic idea is to consider one component of the variable \( v \) as a target and the rest as nuisance. In all cases, we provide accurate closed form results under realistic assumptions. The results are illustrated by one numerical example for the BTL model.

math.OC

A unified theory of the high-dimensional Laplace approximation with application to Bayesian inverse problems

The Laplace approximation (LA) to posteriors is a ubiquitous tool to simplify Bayesian computation, particularly in the high-dimensional settings arising in Bayesian inverse problems. Precisely quantifying the LA accuracy is a challenging problem in the high-dimensional regime. We develop a theory of the LA accuracy to high-dimensional posteriors which both subsumes and unifies a number of results in the literature. The primary advantage of our theory is that we introduce a new degree of flexibility, which can be used to obtain problem-specific upper bounds which are much tighter than previous "rigid" bounds. We demonstrate the theory in a prototypical example of a Bayesian inverse problem, in which this flexibility enables us to improve on prior bounds by an order of magnitude. Our optimized bounds in this setting are dimension-free, and therefore valid in arbitrarily high dimensions.

math.ST

Finite sample expansions and risk bounds in high-dimensional SLS models

This note extends the results of classical parametric statistics like Fisher and Wilks theorem to modern setups with a high or infinite parameter dimension, limited sample size, and possible model misspecification. We consider a special class of stochastically linear smooth (SLS) models satisfying three major conditions: the stochastic component of the log-likelihood is linear in the model parameter and the expected log-likelihood is a smooth and concave function. For the penalized maximum likelihood estimators (pMLE), we establish three types of results: (1) concentration in a small vicinity of the ``truth''; (2) Fisher and Wilks expansions; (3) risk bounds. In all results, the remainder is given explicitly and can be evaluated in terms of the effective sample size and effective parameter dimension which allows us to identify the so-called \emph{critical parameter dimension}. The results are also dimension and coordinate-free. The obtained finite sample expansions are of special interest because they can be used not only for obtaining the risk bounds but also for inference, studying the asymptotic distribution, analysis of resampling procedures, etc. The main tool for all these expansions is the so-called ``basic lemma'' about linearly perturbed optimization. Despite their generality, all the presented bounds are nearly sharp and the classical asymptotic results can be obtained as simple corollaries. Our results indicate that the use of advanced fourth-order expansions allows to relax the critical dimension condition $ \mathbb{p}^{3} \ll n $ from Spokoiny (2023a) to $ \mathbb{p}^{3/2} \ll n $. Examples for classical models like logistic regression, log-density and precision matrix estimation illustrate the applicability of general results.

math.ST

Sharp bounds in perturbed smooth optimization

This paper studies the problem of perturbed convex and smooth optimization. The main results describe how the solution and the value of the problem change if the objective function is perturbed. Examples include linear, quadratic, and smooth additive perturbations. Such problems naturally arise in statistics and machine learning, stochastic optimization, stability and robustness analysis, inverse problems, optimal control, etc. The results provide accurate expansions for the difference between the solution of the original problem and its perturbed counterpart with an explicit error term.

math.OC

Semiparametric plug-in estimation, sup-norm risk bounds, marginal optimization, and inference in BTL model

The recent paper \cite{GSZ2023} on estimation and inference for top-ranking problem in Bradley-Terry-Lice (BTL) model presented a surprising result: component-wise estimation and inference can be done under much weaker conditions on the number of comparison then it is required for the full dimensional estimation. The present paper revisits this finding from completely different viewpoint. Namely, we show how a theoretical study of \emph{estimation in sup-norm} can be reduced to the analysis of \emph{plug-in semiparametric estimation}. For the latter, we adopt and extend the general approach from \cite{Sp2024} to high-dimensional estimation and inference. The main tool of the analysis is a theory of \emph{perturbed marginal optimization} when an objective function depends on a low-dimensional target parameter along with a high-dimensional nuisance parameter. A particular focus of the study is the critical dimension condition. Full-dimensional estimation requires in general the condition \( \mathbbmsl{N} \gg \mathbb{p} \) between the effective parameter dimension \( \mathbb{p} \) and the effective sample size \( \mathbbmsl{N} \) corresponding to the smallest eigenvalue of the Fisher information matrix \( \mathbbmsl{F} \). Inference on the estimated parameter is even more demanding: the condition \( \mathbbmsl{N} \gg \mathbb{p}^{2} \) cannot be generally avoided; see \cite{Sp2024}. However, for the sup-norm estimation, the critical dimension condition can be reduced to \( \mathbbmsl{N} \geq C \log p \).

math.ST

Estimation and inference in error-in-operator model

Many statistical problems can be reduced to a linear inverse problem in which only a noisy version of the operator is available. Particular examples include random design regression, deconvolution problem, instrumental variable regression, functional data analysis, error-in-variable regression, drift estimation in stochastic diffusion, and many others. The pragmatic plug-in approach can be well justified in the classical asymptotic setup with a growing sample size. However, recent developments in high dimensional inference reveal some new features of this problem. In high dimensional linear regression with a random design, the plug-in approach is questionable but the use of a simple ridge penalization yields a benign overfitting phenomenon; see \cite{baLoLu2020}, \cite{ChMo2022}, \cite{NoPuSp2024}. This paper revisits the general Error-in-Operator problem for finite samples and high dimension of the source and image spaces. A particular focus is on the choice of a proper regularization. We show that a simple ridge penalty (Tikhonov regularization) works properly in the case when the operator is more regular than the signal. In the opposite case, some model reduction technique like spectral truncation should be applied.

math.ST

Dimension-free bounds in high-dimensional linear regression via error-in-operator approach

We consider a problem of high-dimensional linear regression with random design. We suggest a novel approach referred to as error-in-operator which does not estimate the design covariance $Σ$ directly but incorporates it into empirical risk minimization. We provide an expansion of the excess prediction risk and derive non-asymptotic dimension-free bounds on the leading term and the remainder. This helps us to show that auxiliary variables do not increase the effective dimension of the problem, provided that parameters of the procedure are tuned properly. We also discuss computational aspects of our method and illustrate its performance with numerical experiments.

math.ST

Estimation and inference for Deep Neuronal Networks

Nonlinear regression problem is one of the most popular and important statistical tasks. The first methods like least squares estimation go back to Gauss and Legendre. Recent models and developments in statistics and machine learning like Deep Neuronal Networks (DNN) or nonlinear PDE stimulate new research in this direction which has to address the important issues and challenges of modern statistical inference such as huge complexity and parameter dimension of the model, limited sample size, lack of convexity and identifiability, among many others. Classical results of nonparametric statistics in terms of rate of convergence do not really address the mentioned issues. This paper offers a general approach to studying a nonlinear regression problem based on the notion of effective dimension. First, a special case of models with stochastically linear structure (SLS) is studied. The results provide finite sample expansions for the loss of the penalized maximum likelihood estimation (MLE). The leading term of such expansions as well as the corresponding remainder are given via the effective dimension and the effective sample size. The obtained expansions can be used to obtain sharp risk bounds and for statistical inference. Despite generality, all the presented bounds are nearly sharp and the classical asymptotic results can be obtained as simple corollaries. Although the basic SLS assumptions are not fulfilled for nonlinear smooth regression, we explain how the stochastic linearity can be achieved by extending the parameter space. The obtained general results are specified to nonlinear smooth regression and to a DNN with one hidden layer.

math.ST

Sharper dimension-free bounds on the Frobenius distance between sample covariance and its expectation

We study properties of a sample covariance estimate $\widehat Σ$ given a finite sample of $n$ i.i.d. centered random elements in $\R^d$ with the covariance matrix $Σ$. We derive dimension-free bounds on the squared Frobenius norm of $(\widehatΣ- Σ)$ under reasonable assumptions. For instance, we show that $\smash{\|\widehatΣ- Σ\|_{\rm F}^2}$ differs from its expectation by at most $\smash{\mathcal O({\rm{Tr}}(Σ^2) / n)}$ with overwhelming probability, which is a significant improvement over the existing results. This allows us to establish the concentration phenomenon for the squared Frobenius distance between the covariance and its empirical counterpart in the case of moderately large effective rank of $Σ$.

math.PR

Concentration of a high dimensional sub-gaussian vector

This note describes the concentration phenomenon for a high dimensional sub-gaussian vector \( X \). In the Gaussian case, for any linear operator \( Q \), it holds \( P\bigl( \| Q X \|^{2} - tr (B) > 2 \sqrt{x\, tr(B^{2})} + 2 \| B \| x \bigr) \leq e^{-x} \) and \( P\bigl( \| Q X \|^{2} - tr (B) < - 2 \sqrt{x \, tr(B^{2})} \bigr) \leq e^{-x} \) with \( B = Q \, Var(X) Q^{T} \); see \cite{laurentmassart2000}. This implies concentration of the squared norm \( \| Q X \|^{2} \) around its expectation \( E \| Q X \|^{2} = tr (B) \) provided that \( tr(B^2)/\| B \|^2 \) is sufficiently large. An extension of this result to a non-gaussian case is a nontrivial task even under sub-gaussian behavior of \( X \), especially if the entries of \( X \) cannot be assumed independent and Hanson-Wright type bounds do not apply. The results of this paper extend the Gaussian deviation bounds and support the concentration phenomenon for \( \| Q X \|^{2} \) using recent advances in Laplace approximation from \cite{SpLaplace2022} and \cite{katsevich2023tight}. The results are illustrated by the case when \( X \) is an i.i.d. sum.

math.PR

Reconstruction of manifold embeddings into Euclidean spaces via intrinsic distances

We consider the problem of reconstructing an embedding of a compact connected Riemannian manifold in a Euclidean space up to an almost isometry, given the information on intrinsic distances between points from its ``sufficiently large'' subset. This is one of the classical manifold learning problems. It happens that the most popular methods to deal with such a problem, with a long history in data science, namely, the classical Multidimensional scaling (MDS) and the Maximum variance unfolding (MVU), actually miss the point and may provide results very far from an isometry; moreover, they may even give no bi-Lipshitz embedding. We will provide an easy variational formulation of this problem, which leads to an algorithm always providing an almost isometric embedding with the distortion of original distances as small as desired (the parameter regulating the upper bound for the desired distortion is an input parameter of this algorithm).

math.OC

Deviation bounds for the norm of a random vector under exponential moment conditions with applications

Hanson-Wright inequality provides a powerful tool for bounding the norm $|ξ|$ of a centered stochastic vector $ξ$ with sub-gaussian behavior. This paper extends the bounds to the case when $ξ$ only has bounded exponential moments of the form $\log E \exp \langle V^{-1} ξ,u \rangle \leq |u|^2/2$, where $V^2 \geq \mathrm{Var}(ξ)$ and $|u| \leq g$ for some fixed $g$. For a linear mapping $Q$, we present an upper quantile function $z_{c}(B,x)$ ensuring $P(| Q ξ| > z_{c}(B,x)) \leq 3 e^{-x}$ with $B = Q \, V^2 Q^{T}$. The obtained results exhibit a phase transition effect: with a value $x_{c}$ depending on $g$ and $B$, for $x \leq x_{c}$, the function $z_{c}(B,x)$ replicates the case of a Gaussian vector $ξ$, that is, $z_{c}^2 (B,x) = {\rm tr}(B) + 2 \sqrt{x {\rm tr}(B^2)} + 2 x |B|$. For $x > x_{c}$, the function $z_{c}(B,x)$ grows linearly in $x$. The results are specified to the case of Bernoulli vector sums and to covariance estimation in Frobenius norm.

math.PR

Mixed Laplace approximation for marginal posterior and Bayesian inference in error-in-operator model

Laplace approximation is a very useful tool in Bayesian inference and it claims a nearly Gaussian behavior of the posterior. \cite{SpLaplace2022} established some rather accurate finite sample results about the quality of Laplace approximation in terms of the so called effective dimension $p$ under the critical dimension constraint $p^{3} \ll n$. However, this condition can be too restrictive for many applications like error-in-operator problem or Deep Neuronal Networks. This paper addresses the question whether the dimensionality condition can be relaxed and the accuracy of approximation can be improved if the target of estimation is low dimensional while the nuisance parameter is high or infinite dimensional. Under mild conditions, the marginal posterior can be approximated by a Gaussian mixture and the accuracy of the approximation only depends on the target dimension. Under the condition $p^{2} \ll n$ or in some special situation like semi-orthogonality, the Gaussian mixture can be replaced by one Gaussian distribution leading to a classical Laplace result. The second result greatly benefits from the recent advances in Gaussian comparison from \cite{GNSUl2017}. The results are illustrated and specified for the case of error-in-operator model.

math.ST

Accelerated gradient methods with absolute and relative noise in the gradient

In this paper, we investigate accelerated first-order methods for smooth convex optimization problems under inexact information on the gradient of the objective. The noise in the gradient is considered to be additive with two possibilities: absolute noise bounded by a constant, and relative noise proportional to the norm of the gradient. We investigate the accumulation of the errors in the convex and strongly convex settings with the main difference with most of the previous works being that the feasible set can be unbounded. The key to the latter is to prove a bound on the trajectory of the algorithm. We also give a stopping criterion for the algorithm and consider extensions to the cases of stochastic optimization and composite nonsmooth problems.

math.OC

Finite samples inference and critical dimension for stochastically linear models

The aim of this note is to state a couple of general results about the properties of the penalized maximum likelihood estimators (pMLE) and of the posterior distribution for parametric models in a non-asymptotic setup and for possibly large or even infinite parameter dimension. We consider a special class of stochastically linear smooth (SLS) models satisfying two major conditions: the stochastic component of the log-likelihood is linear in the model parameter, while the expected log-likelihood is a smooth function. The main results simplify a lot if the expected log-likelihood is concave. For the pMLE, we establish a number of finite sample bounds about its concentration and large deviations as well as the Fisher and Wilks expansion. The later results extend the classical asymptotic Fisher and Wilks Theorems about the MLE to the non-asymptotic setup with large parameter dimension which can depend on the sample size. For the posterior distribution, our main result states a Gaussian approximation of the posterior which can be viewed as a finite sample analog of the prominent Bernstein--von Mises Theorem. In all bounds, the remainder is given explicitly and can be evaluated in terms of the effective sample size and effective parameter dimension. The results are dimension and coordinate free. In spite of generality, all the presented bounds are nearly sharp and the classical asymptotic results can be obtained as simple corollaries. Interesting cases of logit regression and of estimation of a log-density with smooth or truncation priors are used to specify the results and to explain the main notions.

math.ST

Adaptive Weights Community Detection

Due to the technological progress of the last decades, Community Detection has become a major topic in machine learning. However, there is still a huge gap between practical and theoretical results, as theoretically optimal procedures often lack a feasible implementation and vice versa. This paper aims to close this gap and presents a novel algorithm that is both numerically and statistically efficient. Our procedure uses a test of homogeneity to compute adaptive weights describing local communities. The approach was inspired by the Adaptive Weights Community Detection (AWCD) algorithm by Adamyan et al. (2019). This algorithm delivered some promising results on artificial and real-life data, but our theoretical analysis reveals its performance to be suboptimal on a stochastic block model. In particular, the involved estimators are biased and the procedure does not work for sparse graphs. We propose significant modifications, addressing both shortcomings and achieving a nearly optimal rate of strong consistency on the stochastic block model. Our theoretical results are illustrated and validated by numerical experiments.

math.ST