SearcharxivSearch

arXiv subjects

Iqraa Meah

Publications and source records attributed to Iqraa Meah.

4 recordsLinked to original sources

Improved null proportion estimators for multiple discrete tests with plug-in FDR control

It is well known that the performance of the Benjamini and Hochberg (BH) procedure can be improved by incorporating estimators of the number or proportion of null hypotheses to yield an adaptive BH procedure which still controls FDR. Several such plug in estimators have been proposed. For some of these, such as Storey's estimator, plug in FDR control has been established, while for some others, such as the Pounds and Cheng estimator, some gaps remain to be closed. These developments have largely focused on the case of continuous test statistics, where null p values follow the uniform distribution. In the discrete setting, although these estimators continue to provide plug in FDR control, they become overly conservative, leading to inefficient procedures. In this paper, a general class of estimators that encompasses the classical Storey and Pounds and Cheng estimators is introduced. Alongside, several generic strategies to mitigate conservativeness in the discrete setting are proposed by incorporating information about the null distribution functions. These strategies provably yield less conservative estimates while maintaining valid FDR control, and the resulting performance gains are illustrated on both real and simulated data. As a byproduct of a more general result, plug in FDR control for the Pounds and Cheng estimator in the continuous case is also established.

stat.ME

False discovery proportion envelopes with m-consistency

We provide new non-asymptotic false discovery proportion (FDP) confidence envelopes in several multiple testing settings relevant for modern high dimensional-data methods. We revisit the multiple testing scenarios considered in the recent work of Katsevich and Ramdas (2020): top-$k$, preordered (including knockoffs), online. Our emphasis is on obtaining FDP confidence bounds that both have non-asymptotic coverage and are asymptotically accurate in a specific sense, as the number $m$ of tested hypotheses grows. Namely, we introduce and study the property (which we call $m$-consistency) that the confidence bound converges to or below the desired level $\alpha$ when applied to a specific reference $\alpha$-level false discovery rate (FDR) controlling procedure. In this perspective, we derive new bounds that provide improvements over existing ones, both theoretically and practically, and are suitable for situations where at least a moderate number of rejections is expected. These improvements are illustrated with numerical experiments and real data examples. In particular, the improvement is significant in the knockoffs setting, which shows the impact of the method for a practical use. As side results, we introduce a new confidence envelope for the empirical cumulative distribution function of i.i.d. uniform variables, and we provide new power results in sparse cases, both being of independent interest.

math.ST

Online multiple testing with super-uniformity reward

Valid online inference is an important problem in contemporary multiple testing research,to which various solutions have been proposed recently. It is well-known that these existing methods can suffer from a significant loss of power if the null $p$-values are conservative. In this work, we extend the previously introduced methodology to obtain more powerful procedures for the case of super-uniformly distributed $p$-values. These types of $p$-values arise in important settings, e.g. when discrete hypothesis tests are performed or when the $p$-values are weighted. To this end, we introduce the method of super-uniformity reward (SUR) that incorporates information about the individual null cumulative distribution functions. Our approach yields several new 'rewarded' procedures that offer uniform power improvements over known procedures and come with mathematical guarantees for controlling online error criteria based either on the family-wise error rate (FWER) or the marginal false discovery rate (mFDR). We illustrate the benefit of super-uniform rewarding in real-data analyses and simulation studies. While discrete tests serve as our leading example, we also show how our method can be applied to weighted $p$-values.

stat.ME

Differentially Private Federated Learning for Cancer Prediction

Since 2014, the NIH funded iDASH (integrating Data for Analysis, Anonymization, SHaring) National Center for Biomedical Computing has hosted yearly competitions on the topic of private computing for genomic data. For one track of the 2020 iteration of this competition, participants were challenged to produce an approach to federated learning (FL) training of genomic cancer prediction models using differential privacy (DP), with submissions ranked according to held-out test accuracy for a given set of DP budgets. More precisely, in this track, we are tasked with training a supervised model for the prediction of breast cancer occurrence from genomic data split between two virtual centers while ensuring data privacy with respect to model transfer via DP. In this article, we present our 3rd place submission to this competition. During the competition, we encountered two main challenges discussed in this article: i) ensuring correctness of the privacy budget evaluation and ii) achieving an acceptable trade-off between prediction performance and privacy budget.

stat.ML