SearcharxivSearch

arXiv subjects

Guillermo Durand

Publications and source records attributed to Guillermo Durand.

7 recordsLinked to original sources

Confidence envelopes for the false discoveries with heterogeneous data

In the context of selective inference, confidence envelopes for the false discoveries allow the user to select any subset of null hypotheses while having a statistical guarantee on the number of false discoveries in the selected set. Many constructions of such envelopes have been proposed recently, using local test families (Genovese and Wasserman, 2006; Goeman and Solari, 2011), paths (Katsevich and Ramdas, 2020) or interpolation (Blanchard et al., 2020a). All those methods have in common that they have been well-studied for the homogeneous case where all p-values under the null have a uniform distribution over [0, 1]. However, in many applications the data are heterogeneous and discrete, hence the p-values have heterogeneous, discrete distributions, and the previous constructions may incur a loss of power, in the sense that they over-estimate the number of false discoveries. In this paper, we bridge the previous constructions under the homogeneous case with new tools. We also apply these tools to propose several confidence envelopes based on tools tailored for heterogeneous data, like the Bretagnolle inequality, or a new variant of the Simes inequality. We compare these new envelopes to their homogeneous counterparts on simulated data.

math.ST

Fast confidence bounds for the false discovery proportion over a path of hypotheses

This paper presents a new algorithm (and an additional trick) that allows to compute fastly an entire curve of post hoc bounds for the False Discovery Proportion when the underlying bound $V^*\_{\mathfrak{R}}$ construction is based on a reference family $\mathfrak{R}$ with a forest structure {\`a} la Durand et al. (2020). By an entire curve, we mean the values $V^*\_{\mathfrak{R}}(S\_1),\dotsc,V^*\_{\mathfrak{R}}(S\_m)$ computed on a path of increasing selection sets $S\_1\subsetneq\dotsb\subsetneq S\_m$, $|S\_t|=t$. The new algorithm leverages the fact that going from $S\_t$ to $S\_{t+1}$ is done by adding only one hypothesis. Compared to a more naive approach, the new algorithm has a complexity in $O(|\mathcal K|m)$ instead of $O(|\mathcal K|m^2)$, where $|\mathcal K|$ is the cardinality of the family.

math.ST

FDR control and FDP bounds for conformal link prediction

In Marandon (2023), the author introduces a procedure to detect true edges from a partially observed graph using a conformal prediction fashion: first computing scores from a trained function, deriving conformal p-values from them and finally applying a multiple testing procedure. In this paper, we prove that the resulting procedure indeed controls the FDR, and we also derive uniform FDP bounds, thanks to an exchangeability argument and the previous work of Marandon et al. (2022).

math.ST

DiscreteFDR: An R package for controlling the false discovery rate for discrete test statistics

The simultaneous analysis of many statistical tests is ubiquitous in applications. Perhaps the most popular error rate used for avoiding type one error inflation is the false discovery rate (FDR). However, most theoretical and software development for FDR control has focused on the case of continuous test statistics. For discrete data, methods that provide proven FDR control and good performance have been proposed only recently. The R package DiscreteFDR provides an implementation of these methods. For particular commonly used discrete tests such as Fisher's exact test, it can be applied as an off-the-shelf tool by taking only the raw data as input. It can also be used for any arbitrary discrete test statistics by using some additional information on the distribution of these statistics. The paper reviews the statistical methods in a non-technical way, provides a detailed description of the implementation in DiscreteFDR and presents some sample code and analyses.

stat.CO

Adaptive p-value weighting with power optimality

Weighting the p-values is a well-established strategy that improves the power of multiple testing procedures while dealing with heterogeneous data. However, how to achieve this task in an optimal way is rarely considered in the literature. This paper contributes to fill the gap in the case of group-structured null hypotheses, by introducing a new class of procedures named ADDOW (for Adaptive Data Driven Optimal Weighting) that adapts both to the alternative distribution and to the proportion of true null hypotheses. We prove the asymptotical FDR control and power optimality among all weighted procedures of ADDOW, which shows that it dominates all existing procedures in that framework. Some numerical experiments show that the proposed method preserves its optimal properties in the finite sample setting when the number of tests is moderately large.

math.ST

Post hoc false positive control for spatially structured hypotheses

In a high dimensional multiple testing framework, we present new confidence bounds on the false positives contained in subsets S of selected null hypotheses. The coverage probability holds simultaneously over all subsets S, which means that the obtained confidence bounds are post hoc. Therefore, S can be chosen arbitrarily, possibly by using the data set several times. We focus in this paper specifically on the case where the null hypotheses are spatially structured. Our method is based on recent advances in post hoc inference and particularly on the general methodology of Blanchard et al. (2017); we build confidence bounds for some pre-specified forest-structured subsets {R k , k $\in$ K}, called the reference family, and then we deduce a bound for any subset S by interpolation. The proposed bounds are shown to improve substantially previous ones when the signal is locally structured. Our findings are supported both by theoretical results and numerical experiments. Moreover, we show that our bound can be obtained by a low-complexity algorithm, which makes our approach completely operational for a practical use. The proposed bounds are implemented in the open-source R package sansSouci.

math.ST

Improving the Benjamini-Hochberg Procedure for Discrete Tests

To find interesting items in genome-wide association studies or next generation sequencing data, a crucial point is to design powerful false discovery rate (FDR) controlling procedures that suitably combine discrete tests (typically binomial or Fisher tests). In particular, recent research has been striving for appropriate modifications of the classical Benjamini-Hochberg (BH) step-up procedure that accommodate discreteness. However, despite an important number of attempts, these procedures did not come with theoretical guarantees. The present paper contributes to fill the gap: it presents new modifications of the BH procedure that incorporate the discrete structure of the data and provably control the FDR for any fixed number of null hypotheses (under independence). Markedly, our FDR controlling methodology allows to incorporate simultaneously the discreteness and the quantity of signal of the data (corresponding therefore to a so-called $π\_0$-adaptive procedure). The power advantage of the new methods is demonstrated in a numerical experiment and for some appropriate real data sets.

math.ST