SearcharxivSearch

arXiv subjects

Eitan Greenshtein

Publications and source records attributed to Eitan Greenshtein.

10 recordsLinked to original sources

Empirical Bayes Estimation of the Mean of a Function of the Latent Variable with Applications to the Treatment of Nonresponse

We consider the estimation of linear functionals of the mixing distribution in a nonparametric empirical Bayes framework. Our main interest is in situations in which the mixing distribution is only partially identifiable, as may arise in complex sampling situations with nonresponse. We argue that estimating the functional by applying it to the semiparametric maximum likelihood estimator of the mixing distribution is an efficient tool, even when the maximum likelihood estimator is not unique.

math.ST

Consistent Empirical Bayes Estimation of the Mean of a Mixing Distribution with Applications to Treatment of Nonresponse

We consider a Nonparametric Empirical Bayes (NPEB) framework. Let $Y_i$ be random variables, $Y_i \sim f(y|θ_i)$, $i=1,...,n$, where $θ_i \sim G$, and $θ_i \in Θ$ are independent. The variables $Y_i $ are conditionally independent given $θ_i, \; i=1,...,n$. The mixing distribution $G$ is unknown and assumed to belong to a nonparametric class $\{G \}$. Let $η(θ)$ be a function of $θ$. We address the problem of consistently estimating $E_G η(θ) \equiv η_G$. This problem becomes particularly challenging when $G$ cannot be consistently estimated from the observed data. We motivate this problem, especially in contexts involving nonresponse and missing data. For such cases, a consistent estimation method is suggested and its performance is demonstrated through simulations.

math.ST

Consistent Empirical Bayes estimation of the mean of a mixing distribution without identifiability assumption. With applications to treatment of non-response

{\bf Abstract} Consider a Non-Parametric Empirical Bayes (NPEB) setup. We observe $Y_i, \sim f(y|θ_i)$, $θ_i \in Θ$ independent, where $θ_i \sim G$ are independent $i=1,...,n$. The mixing distribution $G$ is unknown $G \in \{G\}$ with no parametric assumptions about the class $\{G \}$. The common NPEB task is to estimate $θ_i, \; i=1,...,n$. Conditions that imply 'optimality' of such NPEB estimators typically require identifiability of $G$ based on $Y_1,...,Y_n$. We consider the task of estimating $E_G θ$. We show that `often' consistent estimation of $E_G θ$ is implied without identifiability. We motivate the later task, especially in setups with non-response and missing data. We demonstrate consistency in simulations.

math.ST

Generalized maximum likelihood estimation of the mean of parameters of mixtures, with applications to sampling

Let $f(y|θ), \; θ\in Ω$ be a parametric family, $η(θ)$ a given function, and $G$ an unknown mixing distribution. It is desired to estimate $E_G (η(θ))\equiv η_G$ based on independent observations $Y_1,...,Y_n$, where $Y_i \sim f(y|θ_i)$, and $θ_i \sim G$ are iid. We explore the Generalized Maximum Likelihood Estimators (GMLE) for this problem. Some basic properties and representations of those estimators are shown. In particular we suggest a new perspective, of the weak convergence result by Kiefer and Wolfowitz (1956), with implications to a corresponding setup in which $θ_1,...,θ_n$ are {\it fixed} parameters. We also relate the above problem, of estimating $η_G$, to non-parametric empirical Bayes estimation under a squared loss. Applications of GMLE to sampling problems are presented. The performance of the GMLE is demonstrated both in simulations and through a real data example.

math.ST

Generalized Maximum Likelihood Estimators and their applications to stratified sampling and post-stratification with many unobserved strata

Consider the problem of estimating a weighted average of the means of $n$ strata, based on a random sample with realized $K_i$ observations from stratum $i, \; i=1,...,n$. This task is non-trivial in cases where for a significant portion of the strata the corresponding $K_i=0$. Such a situation may happen in post-stratification, when it is desired to have a very fine sftratification. A fine stratification could be desired in order that assumptions, or, approximations, like Missing At Random conditional on strata, will be appealing. A fine stratification could also be desired in observational studies, when it is desired to estimate average treatment effect, by averaging the effects in small and homogenous strata. Our approach is based on applying Generalized Maximum Likelihood Estimators (GMLE), and ideas that are related to Non-Parametric Empirical Bayes, in order to estimate the means of strata $i$ with corresponding $K_i=0$. There are no assumptions about a relation between the means of the unobserved strata (i.e., with $K_i=0$) and those of the observed strata. The performance of our approach is demonstrated both in simulations and on a real data set. Some consistency and asymptotic results are also presented. In addition, related basic results about GMLE estimation of the mean of mixtures of exponential families are provided.

math.ST

Deconvolution, convex optimization, non-parametric empirical Bayes and treatment of non-response

Let $(Y_i,θ_i)$, $i=1,...,n$, be independent random vectors distributed like $(Y,θ) \sim G^*$, where the marginal distribution of $θ$ is completely unknown, and the conditional distribution of $Y$ conditional on $θ$ is known. It is desired to estimate the marginal distribution of $θ$ under $G^*$, as well as functionals of the form $E_{G^*} h(Y,θ)$ for a given $h$, based on the observed $Y_1,...,Y_n$. In this paper we suggest a deconvolution method for the above estimation problems and discuss some of its applications in Empirical Bayes analysis. The method involves a quadratic programming step, which is an elaboration on the formulation and technique in Efron(2013). It is computationally efficient and may handle large data sets, where the popular method, of deconvolution using EM-algorithm, is impractical. The main application that we study is treatment of non-response. Our approach is nonstandard and does not involve missing at random type of assumptions. The method is demonstrated in simulations, as well as in an analysis of a real data set from the Labor force survey in Israel. Other applications including estimation of the risk, and estimation of False Discovery Rates, are also discussed. We also present a method, that involves convex optimization, for constructing confidence intervals for $E_{G^*} h$, under the above setup.

math.ST

Deconvolution with application to estimation of sampling probabilities and the Horvitz-Thompson estimator

We elaborate on a deconvolution method, used to estimate the empirical distribution of unknown parameters, as suggested recently by Efron (2013). It is applied to estimating the empirical distribution of the 'sampling probabilities' of m sampled items. The estimated empirical distribution is used to modify the Horvitz-Thompson estimator. The performance of the modified Horvitz-Thompson estimator is studied in two examples. In one example the sampling probabilities are estimated based on the number of visits until a response was obtained. The other example is based on real data from panel sampling, where in four consecutive months there are corresponding four attempts to interview each member in a panel. The sampling probabilities are estimated based on the number of successful attempts. We also discuss briefly, further applications of deconvolution, including estimation of False discovery rate.

math.ST

Nonparametric empirical Bayes and compound decision approaches to estimation of a high-dimensional vector of normal means

We consider the classical problem of estimating a vector $\boldsμ=(μ_1,...,μ_n)$ based on independent observations $Y_i\sim N(μ_i,1)$, $i=1,...,n$. Suppose $μ_i$, $i=1,...,n$ are independent realizations from a completely unknown $G$. We suggest an easily computed estimator $\hat{\boldsμ}$, such that the ratio of its risk $E(\hat{\boldsμ}-\boldsμ)^2$ with that of the Bayes procedure approaches 1. A related compound decision result is also obtained. Our asymptotics is of a triangular array; that is, we allow the distribution $G$ to depend on $n$. Thus, our theoretical asymptotic results are also meaningful in situations where the vector $\boldsμ$ is sparse and the proportion of zero coordinates approaches 1. We demonstrate the performance of our estimator in simulations, emphasizing sparse setups. In ``moderately-sparse'' situations, our procedure performs very well compared to known procedures tailored for sparse setups. It also adapts well to nonsparse situations.

math.ST

Asymptotic efficiency of simple decisions for the compound decision problem

We consider the compound decision problem of estimating a vector of $n$ parameters, known up to a permutation, corresponding to $n$ independent observations, and discuss the difference between two symmetric classes of estimators. The first and larger class is restricted to the set of all permutation invariant estimators. The second class is restricted further to simple symmetric procedures. That is, estimators such that each parameter is estimated by a function of the corresponding observation alone. We show that under mild conditions, the minimal total squared error risks over these two classes are asymptotically equivalent up to essentially O(1) difference.

math.ST

Best subset selection, persistence in high-dimensional statistical learning and optimization under $l_1$ constraint

Let $(Y,X_1,...,X_m)$ be a random vector. It is desired to predict $Y$ based on $(X_1,...,X_m)$. Examples of prediction methods are regression, classification using logistic regression or separating hyperplanes, and so on. We consider the problem of best subset selection, and study it in the context $m=n^α$, $α>1$, where $n$ is the number of observations. We investigate procedures that are based on empirical risk minimization. It is shown, that in common cases, we should aim to find the best subset among those of size which is of order $o(n/\log(n))$. It is also shown, that in some ``asymptotic sense,'' when assuming a certain sparsity condition, there is no loss in letting $m$ be much larger than $n$, for example, $m=n^α, α>1$. This is in comparison to starting with the ``best'' subset of size smaller than $n$ and regardless of the value of $α$. We then study conditions under which empirical risk minimization subject to $l_1$ constraint yields nearly the best subset. These results extend some recent results obtained by Greenshtein and Ritov. Finally we present a high-dimensional simulation study of a ``boosting type'' classification procedure.

math.ST