SearcharxivSearch

arXiv subjects

Eduard Belitser

Publications and source records attributed to Eduard Belitser.

11 recordsLinked to original sources

Bayesian one- and two-sided inference on the local effective dimension

It is a challenge to manage infinite- or high-dimensional data in situations where storage, transmission, or computation resources are constrained. In the simplest scenario when the data consists of a noisy infinite-dimensional signal, we introduce the notion of local \emph{effective dimension} (i.e., pertinent to the underlying signal), formulate and study the problem of its recovery on the basis of noisy data. This problem can be associated to the problems of adaptive quantization, (lossy) data compression, oracle signal estimation. We apply a Bayesian approach and study frequentists properties of the resulting posterior, a purely frequentist version of the results is also proposed. We derive certain upper and lower bounds results about identifying the local effective dimension which show that only the so called \emph{one-sided inference} on the local effective dimension can be ensured whereas the \emph{two-sided inference}, on the other hand, is in general impossible. We establish the \emph{minimal} conditions under which two-sided inference can be made. Finally, connection to the problem of smoothness estimation for some traditional smoothness scales (Sobolev scales) is considered.

math.ST

Robust oracle estimation and uncertainty quantification for possibly sparse quantiles

A general many quantiles + noise model is studied in the robust formulation (allowing non-normal, non-independent observations), where the identifiability requirement for the noise is formulated in terms of quantiles rather than the traditional zero expectation assumption. We propose a penalization method based on the quantile loss function with appropriately chosen penalty function making inference on possibly sparse high-dimensional quantile vector. We apply a local approach to address the optimality by comparing procedures to the oracle sparsity structure. We establish that the proposed procedure mimics the oracle in the problems of estimation and uncertainty quantification (under the so called EBR condition). Adaptive minimax results over sparsity scale follow from our local results.

math.ST

Uncertainty quantification for robust variable selection and multiple testing

We study the problem of identifying the set of \emph{active} variables, termed in the literature as \emph{variable selection} or \emph{multiple hypothesis testing}, depending on the pursued criteria. For a general \emph{robust setting} of non-normal, possibly dependent observations and a generalized notion of \emph{active set}, we propose a procedure that is used simultaneously for the both tasks, variable selection and multiple testing. The procedure is based on the \emph{risk hull minimization} method, but can also be obtained as a result of an empirical Bayes approach or a penalization strategy. We address its quality via various criteria: the Hamming risk, FDR, FPR, FWER, NDR, FNR,and various \emph{multiple testing risks}, e.g., MTR=FDR+NDR; and discuss a weak optimality of our results. Finally, we introduce and study, for the first time, the \emph{uncertainty quantification problem} in the variable selection and multiple testing context in our robust setting.

math.ST

General framework for projection structures

In the first part, we develop a general framework for projection structures and study several inference problems within this framework. We propose procedures based on data dependent measures (DDM) and make connections with empirical Bayes and penalization methods. The main inference problem is the uncertainty quantification (UQ), but on the way we solve the estimation, DDM-contraction problems, and a weak version of the structure recovery problem. The approach is local in that the quality of the inference procedures is measured by the local quantity, the oracle rate, which is the best trade-off between the approximation error by a projection structure and the complexity of that approximating projection structure. Like in statistical learning settings, we develop distribution-free theory as no particular model is imposed, we only assume certain mild condition on the stochastic part of the projection predictor. We introduce the excessive bias restriction (EBR) under which we establish the local confidence optimality of the constructed confidence ball. The proposed general framework unifies a very broad class of high-dimensional models and structures, interesting and important on their own right. In the second part, we apply the developed theory and demonstrate how the general results deliver a whole avenue of local and global minimax results (many new ones, some known results from the literature are improved) for particular models and structures as consequences, including white noise model and density estimation with smoothness structure, linear regression and dictionary learning with sparsity structures, biclustering and stochastic block models with clustering structure, covariance matrix estimation with banding and sparsity structures, and many others. Various adaptive minimax results over various scales follow also from our local results.

math.ST

Needles and straw in a haystack: robust confidence for possibly sparse sequences

In the general signal+noise model we construct an empirical Bayes posterior which we then use for uncertainty quantification for the unknown, possibly sparse, signal. We introduce a novel excessive bias restriction (EBR) condition, which gives rise to a new slicing of the entire space that is suitable for uncertainty quantification. Under EBR and some mild conditions on the noise, we establish the local (oracle) optimality of the proposed confidence ball. In passing, we also get the local optimal (oracle) results for estimation and posterior contraction problems. Adaptive minimax results (also for the estimation and posterior contraction problems) over various sparsity classes follow from our local results.

math.ST

On coverage and local radial rates of DDM-credible sets

For a general statistical model, we introduce the notion of data dependent measure (DDM) on the model parameter. Typical examples of DDM are the posterior distributions. Like for posteriors, the quality of a DDM is characterized by the contraction rate which we allow to be local, i.e., depending on the parameter. We construct confidence sets as DDM-credible sets and address the issue of optimality of such sets, via a trade-off between its "size" (the local radial rate) and its coverage probability. In the mildly ill-posed inverse signal-in-white-noise model, we construct a DDM as empirical Bayes posterior with respect to a certain prior, and define its (default) credible set. Then we introduce 'excessive bias restriction' (EBR), more general than 'self-similarity' and 'polished tail condition' recently studied in the literature. Under EBR, we establish the confidence optimality of our credible set with some local (oracle) radial rate. We also derive the oracle estimation inequality and the oracle DDM-contraction rate, non-asymptotically and uniformly in $\ell_2$. The obtained local results are more powerful than global: adaptive minimax results for a number of smoothness scales follow as consequence, in particular, the ones considered by Szabo, van der Vaart and van Zanten (2015).

math.ST

Pinsker bound under measurement budget constrain: optimal allocation

In the classical many normal means with different variances, we consider the situation when the observer is allowed to allocate the available measurement budget over the coordinates of the parameter of interest. The benchmark is the minimax linear risk over a set. We solve the problem of optimal allocation of observations under the measurement budget constrain for two types of sets, ellipsoids and hyperrectangles. By elaborating on the two examples of Sobolev ellipsoids and hyperectangles, we demonstrate how re-allocating the measurements in the (sub-)optimal way improves on the standard uniform allocation. In particular, we improve the famous Pinsker (1980) bound.

math.ST

Rate-optimal Bayesian intensity smoothing for inhomogeneous Poisson processes

We apply nonparametric Bayesian methods to study the problem of estimating the intensity function of an inhomogeneous Poisson process. We exhibit a prior on intensities which both leads to a computationally feasible method and enjoys desirable theoretical optimality properties. The prior we use is based on B-spline expansions with free knots, adapted from well-established methods used in regression, for instance. We illustrate its practical use in the Poisson process setting by analyzing count data coming from a call centre. Theoretically we derive a new general theorem on contraction rates for posteriors in the setting of intensity function estimation. Practical choices that have to be made in the construction of our concrete prior, such as choosing the priors on the number and the locations of the spline knots, are based on these theoretical findings. The results assert that when properly constructed, our approach yields a rate-optimal procedure that automatically adapts to the regularity of the unknown intensity function.

math.ST

Online Tracking of a Predictable Drifting Parameter of a Time Series

We propose an online algorithm for tracking a multidimensional time-varying parameter of a time series, which is also allowed to be a predictable process with respect to the underlying time series. The algorithm is driven by a gain function. Under assumptions on the gain, we derive uniform non-asymptotic error bounds on the tracking algorithm in terms of chosen step size for the algorithm and the variation of the parameter of interest. We also outline how appropriate gain functions can be constructed. We give several examples of different variational setups for the parameter process where our result can be applied. The proposed approach covers many frameworks and models (including the classical Robbins-Monro and Kiefer-Wolfowitz procedures) where stochastic approximation algorithms comprise the main inference tool for the data analysis. We treat in some detail a couple of specific models.

math.ST

Adaptive Priors based on Splines with Random Knots

Splines are useful building blocks when constructing priors on nonparametric models indexed by functions. Recently it has been established in the literature that hierarchical priors based on splines with a random number of equally spaced knots and random coefficients in the B-spline basis corresponding to those knots lead, under certain conditions, to adaptive posterior contraction rates, over certain smoothness functional classes. In this paper we extend these results for when the location of the knots is also endowed with a prior. This has already been a common practice in MCMC applications, where the resulting posterior is expected to be more "spatially adaptive", but a theoretical basis in terms of adaptive contraction rates was missing. Under some mild assumptions, we establish a result that provides sufficient conditions for adaptive contraction rates in a range of models.

math.ST

Optimal two-stage procedures for estimating location and size of the maximum of a multivariate regression function

We propose a two-stage procedure for estimating the location $\boldsμ$ and size M of the maximum of a smooth d-variate regression function f. In the first stage, a preliminary estimator of $\boldsμ$ obtained from a standard nonparametric smoothing method is used. At the second stage, we "zoom-in" near the vicinity of the preliminary estimator and make further observations at some design points in that vicinity. We fit an appropriate polynomial regression model to estimate the location and size of the maximum. We establish that, under suitable smoothness conditions and appropriate choice of the zooming, the second stage estimators have better convergence rates than the corresponding first stage estimators of $\boldsμ$ and M. More specifically, for $α$-smooth regression functions, the optimal nonparametric rates $n^{-(α-1)/(2α+d)}$ and $n^{-α/(2α+d)}$ at the first stage can be improved to $n^{-(α-1)/(2α)}$ and $n^{-1/2}$, respectively, for $α>1+\sqrt{1+d/2}$. These rates are optimal in the class of all possible sequential estimators. Interestingly, the two-stage procedure resolves "the curse of the dimensionality" problem to some extent, as the dimension d does not control the second stage convergence rates, provided that the function class is sufficiently smooth. We consider a multi-stage generalization of our procedure that attains the optimal rate for any smoothness level $α>2$ starting with a preliminary estimator with any power-law rate at the first stage.

math.ST