SearcharxivSearch

arXiv subjects

Daniel Yekutieli

Publications and source records attributed to Daniel Yekutieli.

9 recordsLinked to original sources

Hierarchical Bayesian Estimation of Covariance Matrices

We develop a hierarchical Bayesian framework for covariance matrix estimation built on a key observation: while equivariance under the full general linear group GL(p) is well known, it is an extremely restrictive property -- estimators equivariant to GL(p) are limited to scalar multiples of the sample covariance matrix and carry considerably larger risks than shrinkage estimators. By contrast, commonly used shrinkage estimators, including the Haff empirical Bayes estimator, and the Ledoit--Wolf estimators, are all equivariant under the smaller orthogonal group O(p). Exploiting this structure, we establish that the Haar measure Bayes rule in an oracle eigenvalue model is the minimum risk estimator within the class of O(p)-equivariant estimators, and derive oracle Bayes rules for the covariance and precision matrices under the squared Frobenius, Stein, and squared Stein loss functions. These oracle rules serve as theoretical benchmarks that dominate all commonly used estimators. To approximate them when the true eigenvalues are unknown, we introduce a hierarchical Bayes model that places a finite P'olya tree prior on the eigenvalue distribution and uses Gibbs sampling to generate posterior draws, yielding both shrinkage estimates for the eigenvalues and approximations to the oracle Bayes rules. Simulations suggest that the finite P'olya tree prior is able to recover the general form of the distribution of the eigenvalues, and confirm that the resulting estimators closely approach oracle performance, substantially outperforming classical competitors for both covariance and precision matrix estimation.

stat.ME

Nonparametric Shrinkage Estimation in High Dimensional Generalized Linear Models via Polya Trees

Regularization in fitting regression models has been a highly active topic of research in the past few decades, but most of the existing methods are designed for particular situations, e.g. for the case of a sparse coefficient vector. We consider the problem of designing $\textit{universally}$ optimal regularized estimators in a given generalized linear model with fixed effects. First, we propose as a contender the Bayes estimator against an $\textit{ideal}$ prior that assigns equal mass to every permutation of the fixed coefficient vector, thus depending on the true coefficients only through their empirical CDF. We prove some optimality properties of this oracle estimator in both the frequentist and Bayesian frameworks. To compete with the oracle estimator, we posit a hierarchical Bayes model where the individual coefficients are modeled as i.i.d. draws from a common distribution $π$, which is in turn assigned a Polya tree prior to reflect indefiniteness. We demonstrate in examples that the posterior mean of $π$ under the postulated model adapts nonparametrically to the empirical CDF of the true coefficients. Correspondingly, the posterior means of the coefficients themselves are used to mimic the ideal estimator. Numerical experiments show that our method has better estimation and prediction accuracy compared to various parametric and nonparametric alternatives, from relatively standard $L_p$-regularized estimators to modern penalized-likelihood and Bayesian estimators for high dimensional regression.

stat.ME

Distribution-Free Bayesian multivariate predictive inference

We introduce a comprehensive Bayesian multivariate predictive inference framework. The basis for our framework is a hierarchical Bayesian model, that is a mixture of finite Polya trees corresponding to multiple dyadic partitions of the unit cube. Given a sample of observations from an unknown multivariate distribution, the posterior predictive distribution is used to model and generate future observations from the unknown distribution. We illustrate the implementation of our methodology and study its performance on simulated examples. We introduce an algorithm for constructing conformal prediction sets, that provide finite sample probability assurances for future observations, with our Bayesian model.

stat.ME

Exact confidence intervals for the mixing distribution from binomial mixture distribution samples

We present methodology for constructing pointwise confidence intervals for the cumulative distribution function and the quantiles of mixing distributions on the unit interval from binomial mixture distribution samples. No assumptions are made on the shape of the mixing distribution. The confidence intervals are constructed by inverting exact tests of composite null hypotheses regarding the mixing distribution. Our method may be applied to any deconvolution approach that produces test statistics whose distribution is stochastically monotone for stochastic increase of the mixing distribution. We propose a hierarchical Bayes approach, which uses finite Polya Trees for modelling the mixing distribution, that provides stable and accurate deconvolution estimates without the need for additional tuning parameters. Our main technical result establishes the stochastic monotonicity property of the test statistics produced by the hierarchical Bayes approach. Leveraging the need for the stochastic monotonicity property, we explicitly derive the smallest asymptotic confidence intervals that may be constructed using our methodology. Raising the question whether it is possible to construct smaller confidence intervals for the mixing distribution without making parametric assumptions on its shape.

stat.ME

Selective Sign-Determining Multiple Confidence Intervals with FCR Control

Given $m$ unknown parameters with corresponding independent estimators, the Benjamini-Hochberg (BH) procedure can be used to classify the sign of parameters such that the expected proportion of erroneous directional decisions (directional FDR) is controlled at a preset level $q$. More ambitiously, our goal is to construct sign-determining confidence intervals---instead of only classifying the sign---such that the expected proportion of non-covering constructed intervals (FCR) is controlled. We suggest a valid procedure which adjusts a marginal confidence interval in order to construct a maximum number of sign-determining confidence intervals. We propose a new marginal confidence interval, designed specifically for our procedure, which allows to balance a trade-off between power and length of the constructed intervals, and, in fact, often enjoy (almost) the best of both worlds. We apply our methods to detect the sign of correlations in a highly publicized social neuroscience study and, in a second example, to detect the direction of association for SNPs with Type-2 Diabetes in GWAS data. In both examples we compare our procedure to existing methods and obtain encouraging results.

stat.ME

Extracting replicable associations across multiple studies: algorithms for controlling the false discovery rate

Extracting associations that recur across multiple studies while controlling the false discovery rate is a fundamental challenge. Here, we consider an extension of Efron's single-study two-groups model to allow joint analysis of multiple studies. We assume that given a set of p-values obtained from each study, the researcher is interested in associations that recur in at least $k>1$ studies. We propose new algorithms that differ in how the study dependencies are modeled. We compared our new methods and others using various simulated scenarios. The top performing algorithm, SCREEN (Scalable Cluster-based REplicability ENhancement), is our new algorithm that is based on three stages: (1) clustering an estimated correlation network of the studies, (2) learning replicability (e.g., of genes) within clusters, and (3) merging the results across the clusters using dynamic programming. We applied SCREEN to two real datasets and demonstrated that it greatly outperforms the results obtained via standard meta-analysis. First, on a collection of 29 case-control large-scale gene expression cancer studies, we detected a large up-regulated module of genes related to proliferation and cell cycle regulation. These genes are both consistently up-regulated across many cancer studies, and are well connected in known gene networks. Second, on a recent pan-cancer study that examined the expression profiles of patients with or without mutations in the HLA complex, we detected an active module of up-regulated genes that are related to immune responses. Thanks to our ability to quantify the false discovery rate, we detected thrice more genes as compared to the original study. Our module contains most of the genes reported in the original study, and many new ones. Interestingly, the newly discovered genes are needed to establish the connectivity of the module.

stat.ME

Replicability analysis for genome-wide association studies

The paramount importance of replicating associations is well recognized in the genome-wide associaton (GWA) research community, yet methods for assessing replicability of associations are scarce. Published GWA studies often combine separately the results of primary studies and of the follow-up studies. Informally, reporting the two separate meta-analyses, that of the primary studies and follow-up studies, gives a sense of the replicability of the results. We suggest a formal empirical Bayes approach for discovering whether results have been replicated across studies, in which we estimate the optimal rejection region for discovering replicated results. We demonstrate, using realistic simulations, that the average false discovery proportion of our method remains small. We apply our method to six type two diabetes (T2D) GWA studies. Out of 803 SNPs discovered to be associated with T2D using a typical meta-analysis, we discovered 219 SNPs with replicated associations with T2D. We recommend complementing a meta-analysis with a replicability analysis for GWA studies.

stat.ME

Optimal exact tests for composite alternative hypotheses on cross tabulated data

We present methodology for constructing exact significance tests for cross tabulated data for "difficult" composite alternative hypotheses that have no natural test statistic. We construct a test for discovering Simpson's Paradox and a general test for discovering positive dependence between two ordinal variables. Our tests are Bayesian extensions of the likelihood ratio test, they are optimal with respect to the prior distribution, and are also closely related to Bayes factors and Bayesian FDR controlling testing procedures.

stat.ME

Adjusted Bayesian inference for selected parameters

We address the problem of providing inference from a Bayesian perspective for parameters selected after viewing the data. We present a Bayesian framework for providing inference for selected parameters, based on the observation that providing Bayesian inference for selected parameters is a truncated data problem. We show that if the prior for the parameter is non-informative, or if the parameter is a "fixed" unknown constant, then it is necessary to adjust the Bayesian inference for selection. Our second contribution is the introduction of Bayesian False Discovery Rate controlling methodology,which generalizes existing Bayesian FDR methods that are only defined in the two-group mixture model.We illustrate our results by applying them to simulated data and data froma microarray experiment.

stat.CO