SearcharxivSearch

arXiv subjects

Sebastian Döhler

Publications and source records attributed to Sebastian Döhler.

12 recordsLinked to original sources

Confidence envelopes for the false discoveries with heterogeneous data

In the context of selective inference, confidence envelopes for the false discoveries allow the user to select any subset of null hypotheses while having a statistical guarantee on the number of false discoveries in the selected set. Many constructions of such envelopes have been proposed recently, using local test families (Genovese and Wasserman, 2006; Goeman and Solari, 2011), paths (Katsevich and Ramdas, 2020) or interpolation (Blanchard et al., 2020a). All those methods have in common that they have been well-studied for the homogeneous case where all p-values under the null have a uniform distribution over [0, 1]. However, in many applications the data are heterogeneous and discrete, hence the p-values have heterogeneous, discrete distributions, and the previous constructions may incur a loss of power, in the sense that they over-estimate the number of false discoveries. In this paper, we bridge the previous constructions under the homogeneous case with new tools. We also apply these tools to propose several confidence envelopes based on tools tailored for heterogeneous data, like the Bretagnolle inequality, or a new variant of the Simes inequality. We compare these new envelopes to their homogeneous counterparts on simulated data.

math.ST

POLTER: Policy Trajectory Ensemble Regularization for Unsupervised Reinforcement Learning

The goal of Unsupervised Reinforcement Learning (URL) is to find a reward-agnostic prior policy on a task domain, such that the sample-efficiency on supervised downstream tasks is improved. Although agents initialized with such a prior policy can achieve a significantly higher reward with fewer samples when finetuned on the downstream task, it is still an open question how an optimal pretrained prior policy can be achieved in practice. In this work, we present POLTER (Policy Trajectory Ensemble Regularization) - a general method to regularize the pretraining that can be applied to any URL algorithm and is especially useful on data- and knowledge-based URL algorithms. It utilizes an ensemble of policies that are discovered during pretraining and moves the policy of the URL algorithm closer to its optimal prior. Our method is based on a theoretical framework, and we analyze its practical effects on a white-box benchmark, allowing us to study POLTER with full control. In our main experiments, we evaluate POLTER on the Unsupervised Reinforcement Learning Benchmark (URLB), which consists of 12 tasks in 3 domains. We demonstrate the generality of our approach by improving the performance of a diverse set of data- and knowledge-based URL algorithms by 19% on average and up to 40% in the best case. Under a fair comparison with tuned baselines and tuned POLTER, we establish a new state-of-the-art for model-free methods on the URLB.

cs.LG

Improved null proportion estimators for multiple discrete tests with plug-in FDR control

It is well known that the performance of the Benjamini and Hochberg (BH) procedure can be improved by incorporating estimators of the number or proportion of null hypotheses to yield an adaptive BH procedure which still controls FDR. Several such plug in estimators have been proposed. For some of these, such as Storey's estimator, plug in FDR control has been established, while for some others, such as the Pounds and Cheng estimator, some gaps remain to be closed. These developments have largely focused on the case of continuous test statistics, where null p values follow the uniform distribution. In the discrete setting, although these estimators continue to provide plug in FDR control, they become overly conservative, leading to inefficient procedures. In this paper, a general class of estimators that encompasses the classical Storey and Pounds and Cheng estimators is introduced. Alongside, several generic strategies to mitigate conservativeness in the discrete setting are proposed by incorporating information about the null distribution functions. These strategies provably yield less conservative estimates while maintaining valid FDR control, and the resulting performance gains are illustrated on both real and simulated data. As a byproduct of a more general result, plug in FDR control for the Pounds and Cheng estimator in the continuous case is also established.

stat.ME

Contextualize Me -- The Case for Context in Reinforcement Learning

While Reinforcement Learning ( RL) has made great strides towards solving increasingly complicated problems, many algorithms are still brittle to even slight environmental changes. Contextual Reinforcement Learning (cRL) provides a framework to model such changes in a principled manner, thereby enabling flexible, precise and interpretable task specification and generation. Our goal is to show how the framework of cRL contributes to improving zero-shot generalization in RL through meaningful benchmarks and structured reasoning about generalization tasks. We confirm the insight that optimal behavior in cRL requires context information, as in other related areas of partial observability. To empirically validate this in the cRL framework, we provide various context-extended versions of common RL environments. They are part of the first benchmark library, CARL, designed for generalization based on cRL extensions of popular benchmarks, which we propose as a testbed to further study general agents. We show that in the contextual setting, even simple RL environments become challenging - and that naive solutions are not enough to generalize across complex context spaces.

cs.LG

Online multiple testing with super-uniformity reward

Valid online inference is an important problem in contemporary multiple testing research,to which various solutions have been proposed recently. It is well-known that these existing methods can suffer from a significant loss of power if the null $p$-values are conservative. In this work, we extend the previously introduced methodology to obtain more powerful procedures for the case of super-uniformly distributed $p$-values. These types of $p$-values arise in important settings, e.g. when discrete hypothesis tests are performed or when the $p$-values are weighted. To this end, we introduce the method of super-uniformity reward (SUR) that incorporates information about the individual null cumulative distribution functions. Our approach yields several new 'rewarded' procedures that offer uniform power improvements over known procedures and come with mathematical guarantees for controlling online error criteria based either on the family-wise error rate (FWER) or the marginal false discovery rate (mFDR). We illustrate the benefit of super-uniform rewarding in real-data analyses and simulation studies. While discrete tests serve as our leading example, we also show how our method can be applied to weighted $p$-values.

stat.ME

Improved $q$-values for discrete uniform and homogeneous tests: a comparative study

Large scale discrete uniform and homogeneous $P$-values often arise in applications with multiple testing. For example, this occurs in genome wide association studies whenever a nonparametric one-sample (or two-sample) test is applied throughout the gene loci. In this paper we consider $q$-values for such scenarios based on several existing estimators for the proportion of true null hypothesis, $π_0$, which take the discreteness of the $P$-values into account. The theoretical guarantees of the several approaches with respect to the estimation of $π_0$ and the false discovery rate control are reviewed. The performance of the discrete $q$-values is investigated through intensive Monte Carlo simulations, including location, scale and omnibus nonparametric tests, and possibly dependent $P$-values. The methods are applied to genetic and financial data for illustration purposes too. Since the particular estimator of $π_0$ used to compute the $q$-values may influence the power, relative advantages and disadvantages of the reviewed procedures are discussed. Practical recommendations are given.

stat.ME

Controlling false discovery exceedance for heterogeneous tests

Several classical methods exist for controlling the false discovery exceedance (FDX) for large scale multiple testing problems, among them the Lehmann-Romano procedure ([LR] below) and the Guo-Romano procedure ([GR] below). While these two procedures are the most prominent, they were originally designed for homogeneous test statistics, that is, when the null distribution functions of the $p$-values $F_i$, $1\leq i\leq m$, are all equal. In many applications, however, the data are heterogeneous which leads to heterogeneous null distribution functions. Ignoring this heterogeneity usually induces a conservativeness for the aforementioned procedures. In this paper, we develop three new procedures that incorporate the $F_i$'s, while ensuring the FDX control. The heterogeneous version of [LR], denoted [HLR], is based on the arithmetic average of the $F_i$'s, while the heterogeneous version of [GR], denoted [HGR], is based on the geometric average of the $F_i$'s. We also introduce a procedure [PB], that is based on the Poisson-binomial distribution and that uniformly improves [HLR] and [HGR], at the price of a higher computational complexity. Perhaps surprisingly, this shows that, contrary to the known theory of false discovery rate (FDR) control under heterogeneity, the way to incorporate the $F_i$'s can be particularly simple in the case of FDX control, and does not require any further correction term. The performances of the new proposed procedures are illustrated by real and simulated data in two important heterogeneous settings: first, when the test statistics are continuous but the $p$-values are weighted by some known independent weight vector, e.g., coming from co-data sets; second, when the test statistics are discretely distributed, as is the case for data representing frequencies or counts.

stat.ME

DiscreteFDR: An R package for controlling the false discovery rate for discrete test statistics

The simultaneous analysis of many statistical tests is ubiquitous in applications. Perhaps the most popular error rate used for avoiding type one error inflation is the false discovery rate (FDR). However, most theoretical and software development for FDR control has focused on the case of continuous test statistics. For discrete data, methods that provide proven FDR control and good performance have been proposed only recently. The R package DiscreteFDR provides an implementation of these methods. For particular commonly used discrete tests such as Fisher's exact test, it can be applied as an off-the-shelf tool by taking only the raw data as input. It can also be used for any arbitrary discrete test statistics by using some additional information on the distribution of these statistics. The paper reviews the statistical methods in a non-technical way, provides a detailed description of the implementation in DiscreteFDR and presents some sample code and analyses.

stat.CO

Improving the Benjamini-Hochberg Procedure for Discrete Tests

To find interesting items in genome-wide association studies or next generation sequencing data, a crucial point is to design powerful false discovery rate (FDR) controlling procedures that suitably combine discrete tests (typically binomial or Fisher tests). In particular, recent research has been striving for appropriate modifications of the classical Benjamini-Hochberg (BH) step-up procedure that accommodate discreteness. However, despite an important number of attempts, these procedures did not come with theoretical guarantees. The present paper contributes to fill the gap: it presents new modifications of the BH procedure that incorporate the discrete structure of the data and provably control the FDR for any fixed number of null hypotheses (under independence). Markedly, our FDR controlling methodology allows to incorporate simultaneously the discreteness and the quantity of signal of the data (corresponding therefore to a so-called $π\_0$-adaptive procedure). The power advantage of the new methods is demonstrated in a numerical experiment and for some appropriate real data sets.

math.ST

A discrete modification of the Benjamini-Yekutieli procedure

The Benjamini-Yekutieli procedure is a multiple testing method that controls the false discovery rate under arbitrary dependence of the $p$-values. A modification of this and related procedures is proposed for the case when the test statistics are discrete. It is shown that taking discreteness into account can improve upon known procedures. The performance of this new procedure is evaluated for pharmacovigilance data and in a simulation study.

math.ST

A sufficient criterion for control of generalised error rates in multiple testing

Based on the work of Romano and Shaikh (2006) and Lehmann and Romano (2005) we give a sufficient criterion for controlling generalised error rates for arbitrarily dependent p-values. This criterion is formulated in terms of matrices associated with the corresponding error rates and thus it is possible to view the corresponding critical constants as solutions of sets of certain linear inequalities. This property can in some cases be used to improve the power of existing procedures by finding optimal solutions to an associated linear programming problem.

stat.ME

Validation of credit default probabilities via multiple testing procedures

We apply multiple testing procedures to the validation of estimated default probabilities in credit rating systems. The goal is to identify rating classes for which the probability of default is estimated inaccurately, while still maintaining a predefined level of committing type I errors as measured by the familywise error rate (FWER) and the false discovery rate (FDR). For FWER, we also consider procedures that take possible discreteness of the data resp. test statistics into account. The performance of these methods is illustrated in a simulation setting and for empirical default data.

stat.AP