SearcharxivSearch

arXiv subjects

Asaf Weinstein

Publications and source records attributed to Asaf Weinstein.

14 recordsLinked to original sources

Estimating the local false discovery rate under an unknown symmetric null

This paper is concerned with estimating the local false discovery rate (lfdr) in a two-groups model where the only assumption regarding the null distribution is symmetry about zero. Our motivation comes from the contemporary framework for multiple hypothesis testing, particularly relevant in variable selection problems, which transforms any user-specified scores into statistics whose null distributions are symmetric about zero, whereas enrichment to the right of zero is generally expected for the non-nulls. While modern methods such as the knockoff filter (Barber and Candes; 2015) are able to exploit the null property for controlling the false discovery rate (FDR), an arguably more appropriate goal is to target control of the local false discovery rate for the rejected hypotheses, as proposed in Soloff et al. (2024) where the standard two-groups model (known $f_0$ and independence) is analyzed. Here, we take a step in this direction and propose to estimate the lfdr by targeting the surrogate density ratio $f(-w)/f(w)$, for $w>0$, where $f$ is the marginal density in the aforementioned ``stripped-down'' two-groups model. We study several estimators and propose a logistic regression based method with natural cubic spline basis. We also show that any consistent estimator of this surrogate yields asymptotic lfdr control of the multiple testing procedure that thresholds the estimate at the nominal level.

stat.ME

Posterior consistency of P\'olya trees for deconvolution under the linear model

Several recent works have addressed the problem of deconvolution under a linear model, where the goal is to estimate a completely unknown $G_0$ from a vector of noisy observations $\boldsymbol{Y} = X\boldsymbol{\beta} + \boldsymbol{\epsilon}$, assuming the coefficients $\beta_j$ are i.i.d. unobserved realizations from $G_0$. Assuming $G_0$ has a density $g_0$, we study theoretically a Bayesian nonparametric method proposed in Weinstein et al. (2025) that postulates a P\'olya tree prior $\Pi$ on $g_0$ and bases a deconvolution estimate on the posterior distribution $\Pi(\cdot|\boldsymbol{Y})$. Our main result asserts that under the true model (fixed and unknown $g_0$), and under a suitable condition on the minimum eigenvalue of $X^\top X$, the posterior $\Pi(\cdot|\boldsymbol{Y})$ concentrates around $g_0$ in sup-norm. The analysis presented builds on and extends results from Castillo (2017), where posterior consistency of P\'olya trees was proved for density estimation, the simpler problem of estimating $g_0$ when observing the coefficients $\beta_j$ directly.

math.ST

Empirical Bayes Variable Selection with Lasso Statistics in the AMP Framework

The Lasso is one of the most ubiquitous methods for variable selection in high-dimensional linear regression and has been studied extensively under different regimes. In a particular asymptotic setup entailing $n/p\to \text{constant}$, an i.i.d.~Gaussian $X$ matrix and linear sparsity, \citet{su2017false} analyzed the Lasso selection path and presented negative results, showing that maintaining small levels of the false discovery proportion comes at a substantial cost in power. Followup work by \citet{wang2020bridge} used the same framework to study the tradeoff between type I error and power for thresholded-Lasso selection, which ranks the variables based on the magnitude of the Lasso estimate instead of the order of appearance on the Lasso path, and demonstrated that significant improvements are possible if the regularization parameter is chosen appropriately. We take this line of research a step further, seeking an {\em optimal} selection procedure in the AMP framework among procedures that order the variables by some univariate function of the Lasso estimate at a fixed value $\lambda$ of the regularization term. Observing that the model for the Lasso estimates effectively reduces asymptotically to a version of the well-studied two-groups model, we propose an empirical Bayes variable selection procedure based on an estimate of the local false discovery rate. We extend existing results in the AMP framework to obtain exact predictions for the curve describing the asymptotic tradeoff between type I error and power of this procedure. Additionally, we prove that the optimal $\lambda$ is the minimizer of the asymptotic mean squared error, and accordingly propose to use the empirical Bayes procedure with $\lambda$ estimated by cross-validation. The theoretical predictions imply that the gains in power can be substantial, and we confirm this by numerical studies under different settings.

math.ST

On the Minimum Attainable Risk in Permutation Invariant Problems

We introduce a broad class of permutation invariant problems by extending the standard decision theoretic definition to allow also selective inference tasks, where the target is specified only after seeing the data. For any such problem, the minimizer of the risk at $\boldsymbol{\theta}$ among all permutation invariant (equivariant) procedures is shown to be the Bayes rule that posits a uniform prior over all permutations of $\boldsymbol{\theta}$. This gives an explicit form of the greatest lower bound on the risk of any sensible procedure in a wide range of problems. From a practical perspective, approximations to the exact bound are required because of its computational cost. In a specific example of estimating the parameter of a selected population, we prove that our bound coincides asymptotically with the computationally tractable bound attained by the Bayes rule which replaces the uniform prior on all permutations of $\boldsymbol{\theta}$ by the i.i.d. prior with the same marginals. This generalizes results previously known only for the very special case of compound decision problems. The possibility of asymptotically attaining the latter bound by an empirical Bayes rule is discussed.

math.ST

A Power Analysis for Model-X Knockoffs with $\ell_{p}$-Regularized Statistics

Variable selection properties of procedures utilizing penalized-likelihood estimates is a central topic in the study of high dimensional linear regression problems. Existing literature emphasizes the quality of ranking of the variables by such procedures as reflected in the receiver operating characteristic curve or in prediction performance. Specifically, recent works have harnessed modern theory of approximate message-passing (AMP) to obtain, in a particular setting, exact asymptotic predictions of the type I-type II error tradeoff for selection procedures that rely on $\ell_{p}$-regularized estimators. In practice, effective ranking by itself is often not sufficient because some calibration for Type I error is required. In this work we study theoretically the power of selection procedures that similarly rank the features by the size of an $\ell_{p}$-regularized estimator, but further use Model-X knockoffs to control the false discovery rate in the realistic situation where no prior information about the signal is available. In analyzing the power of the resulting procedure, we extend existing results in AMP theory to handle the pairing between original variables and their knockoffs. This is used to derive exact asymptotic predictions for power. We apply the general results to compare the power of the knockoffs versions of Lasso and thresholded-Lasso selection, and demonstrate that in the i.i.d. covariate setting under consideration, tuning by cross-validation on the augmented design matrix is nearly optimal. We further demonstrate how the techniques allow to analyze also the Type S error, and a corresponding notion of power, when selections are supplemented with a decision on the sign of the coefficient.

math.ST

Integrative Methods for Post-Selection Inference Under Convex Constraints

Inference after model selection has been an active research topic in the past few years, with numerous works offering different approaches to addressing the perils of the reuse of data. In particular, major progress has been made recently on large and useful classes of problems by harnessing general theory of hypothesis testing in exponential families, but these methods have their limitations. Perhaps most immediate is the gap between theory and practice: implementing the exact theoretical prescription in realistic situations---for example, when new data arrives and inference needs to be adjusted accordingly---may be a prohibitive task. In this paper we propose a Bayesian framework for carrying out inference after model selection in the linear model. Our framework is very flexible in the sense that it naturally accommodates different models for the data, instead of requiring a case-by-case treatment. At the core of our methods is a new approximation to the exact likelihood conditional on selection, the latter being generally intractable. We prove that, under appropriate conditions, our approximation is asymptotically consistent with the exact truncated likelihood. The advantages of our methods in practical data analysis are demonstrated in simulations and in application to HIV drug-resistance data.

stat.ME

On Optimal Solutions to Compound Statistical Decision Problems

In a compound decision problem, consisting of $n$ statistically independent copies of the same problem to be solved under the sum of the individual losses, any reasonable compound decision rule $δ$ satisfies a natural symmetry property, entailing that $δ(σ(\boldsymbol{y})) = σ(δ(\boldsymbol{y}))$ for any permutation $σ$. We derive the greatest lower bound on the risk of any such decision rule. The classical problem of estimating the mean of a homoscedastic normal vector is used to demonstrate the theory, but important extensions are presented as well in the context of Robbins's original ideas.

math.ST

Nonparametric Shrinkage Estimation in High Dimensional Generalized Linear Models via Polya Trees

Regularization in fitting regression models has been a highly active topic of research in the past few decades, but most of the existing methods are designed for particular situations, e.g. for the case of a sparse coefficient vector. We consider the problem of designing $\textit{universally}$ optimal regularized estimators in a given generalized linear model with fixed effects. First, we propose as a contender the Bayes estimator against an $\textit{ideal}$ prior that assigns equal mass to every permutation of the fixed coefficient vector, thus depending on the true coefficients only through their empirical CDF. We prove some optimality properties of this oracle estimator in both the frequentist and Bayesian frameworks. To compete with the oracle estimator, we posit a hierarchical Bayes model where the individual coefficients are modeled as i.i.d. draws from a common distribution $\pi$, which is in turn assigned a Polya tree prior to reflect indefiniteness. We demonstrate in examples that the posterior mean of $\pi$ under the postulated model adapts nonparametrically to the empirical CDF of the true coefficients. Correspondingly, the posterior means of the coefficients themselves are used to mimic the ideal estimator. Numerical experiments show that our method has better estimation and prediction accuracy compared to various parametric and nonparametric alternatives, from relatively standard $L_p$-regularized estimators to modern penalized-likelihood and Bayesian estimators for high dimensional regression.

stat.ME

Online Control of the False Coverage Rate and False Sign Rate

The false coverage rate (FCR) is the expected ratio of number of constructed confidence intervals (CIs) that fail to cover their respective parameters to the total number of constructed CIs. Procedures for FCR control exist in the offline setting, but none so far have been designed with the online setting in mind. In the online setting, there is an infinite sequence of fixed unknown parameters $θ_t$ ordered by time. At each step, we see independent data that is informative about $θ_t$, and must immediately make a decision whether to report a CI for $θ_t$ or not. If $θ_t$ is selected for coverage, the task is to determine how to construct a CI for $θ_t$ such that $\text{FCR} \leq α$ for any $T\in \mathbb{N}$. A straightforward solution is to construct at each step a $(1-α)$ level conditional CI. In this paper, we present a novel solution to the problem inspired by online false discovery rate (FDR) algorithms, which only requires the statistician to be able to construct a marginal CI at any given level. Apart from the fact that marginal CIs are usually simpler to construct than conditional ones, the marginal procedure has an important qualitative advantage over the conditional solution, namely, it allows selection to be determined by the candidate CI itself. We take advantage of this to offer solutions to some online problems which have not been addressed before. For example, we show that our general CI procedure can be used to devise online sign-classification procedures that control the false sign rate (FSR). In terms of power and length of the constructed CIs, we demonstrate that the two approaches have complementary strengths and weaknesses using simulations. Last, all of our methodology applies equally well to online FCR control for prediction intervals, having particular implications for assumption-free selective conformal inference.

stat.ME

Selective Sign-Determining Multiple Confidence Intervals with FCR Control

Given $m$ unknown parameters with corresponding independent estimators, the Benjamini-Hochberg (BH) procedure can be used to classify the sign of parameters such that the expected proportion of erroneous directional decisions (directional FDR) is controlled at a preset level $q$. More ambitiously, our goal is to construct sign-determining confidence intervals---instead of only classifying the sign---such that the expected proportion of non-covering constructed intervals (FCR) is controlled. We suggest a valid procedure which adjusts a marginal confidence interval in order to construct a maximum number of sign-determining confidence intervals. We propose a new marginal confidence interval, designed specifically for our procedure, which allows to balance a trade-off between power and length of the constructed intervals, and, in fact, often enjoy (almost) the best of both worlds. We apply our methods to detect the sign of correlations in a highly publicized social neuroscience study and, in a second example, to detect the direction of association for SNPs with Type-2 Diabetes in GWAS data. In both examples we compare our procedure to existing methods and obtain encouraging results.

stat.ME

A Power and Prediction Analysis for Knockoffs with Lasso Statistics

Knockoffs is a new framework for controlling the false discovery rate (FDR) in multiple hypothesis testing problems involving complex statistical models. While there has been great emphasis on Type-I error control, Type-II errors have been far less studied. In this paper we analyze the false negative rate or, equivalently, the power of a knockoff procedure associated with the Lasso solution path under an i.i.d. Gaussian design, and find that knockoffs asymptotically achieve close to optimal power with respect to an omniscient oracle. Furthermore, we demonstrate that for sparse signals, performing model selection via knockoff filtering achieves nearly ideal prediction errors as compared to a Lasso oracle equipped with full knowledge of the distribution of the unknown regression coefficients. The i.i.d. Gaussian design is adopted to leverage results concerning the empirical distribution of the Lasso estimates, which makes power calculation possible for both knockoff and oracle procedures.

stat.ME

Group-Linear Empirical Bayes Estimates for a Heteroscedastic Normal Mean

The problem of estimating the mean of a normal vector with known but unequal variances introduces substantial difficulties that impair the adequacy of traditional empirical Bayes estimators. By taking a different approach, that treats the known variances as part of the random observations, we restore symmetry and thus the effectiveness of such methods. We suggest a group-linear empirical Bayes estimator, which collects observations with similar variances and applies a spherically symmetric estimator to each group separately. The proposed estimator is motivated by a new oracle rule which is stronger than the best linear rule, and thus provides a more ambitious benchmark than that considered in previous literature. Our estimator asymptotically achieves the new oracle risk (under appropriate conditions) and at the same time is minimax. The group-linear estimator is particularly advantageous in situations where the true means and observed variances are empirically dependent. To demonstrate the merits of the proposed methods in real applications, we analyze the baseball data used in Brown (2008), where the group-linear methods achieved the prediction error of the best nonparametric estimates that have been applied to the dataset, and significantly lower error than other parametric and semi-parametric empirical Bayes estimators.

stat.ME

Empirical Bayes Estimates for a 2-Way Cross-Classified Additive Model

We develop an empirical Bayes procedure for estimating the cell means in an unbalanced, two-way additive model with fixed effects. We employ a hierarchical model, which reflects exchangeability of the effects within treatment and within block but not necessarily between them, as suggested before by Lindley and Smith (1972). The hyperparameters of this hierarchical model, instead of considered fixed, are to be substituted with data-dependent values in such a way that the point risk of the empirical Bayes estimator is small. Our method chooses the hyperparameters by minimizing an unbiased risk estimate and is shown to be asymptotically optimal for the estimation problem defined above. The usual empirical Best Linear Unbiased Predictor (BLUP) is shown to be substantially different from the proposed method in the unbalanced case and therefore performs sub-optimally. Our estimator is implemented through a computationally tractable algorithm that is scalable to work under large designs. The case of missing cell observations is treated as well. We demonstrate the advantages of our method over the BLUP estimator through simulations and in a real data example, where we estimate average nitrate levels in water sources based on their locations and the time of the day.

stat.ME

Inequalities for the Bayes Risk

Several inequalities are presented which, in part, generalize inequalities by Weinstein and Weiss, giving rise to new lower bounds for the Bayes risk under squared error loss.

cs.IT