SearcharxivSearch

arXiv subjects

Arnold Janssen

Publications and source records attributed to Arnold Janssen.

13 recordsLinked to original sources

Permutation inference in factorial survival designs with the CASANOVA

We propose inference procedures for general nonparametric factorial survival designs with possibly right-censored data. Similar to additive Aalen models, null hypotheses are formulated in terms of cumulative hazards. Thereby, deviations are measured in terms of quadratic forms in Nelson-Aalen-type integrals. Different to existing approaches this allows to work without restrictive model assumptions as proportional hazards. In particular, crossing survival or hazard curves can be detected without a significant loss of power. For a distribution-free application of the method, a permutation strategy is suggested. The resulting procedures' asymptotic validity as well as their consistency are proven and their small sample performances are analyzed in extensive simulations. Their applicability is finally illustrated by analyzing an oncology data set.

stat.ME

Dependence correction of multiple tests with applications to sparsity

The present paper establishes new multiple procedures for simultaneous testing of a large number of hypotheses under dependence. Special attention is devoted to experiments with rare false hypotheses. This sparsity assumption is typically for various genome studies when a portion of remarkable genes should be detected. The aim is to derive tests which control the false discovery rate (FDR) always at finite sample size. The procedures are compared for the set up of dependent and independent $p$-values. It turns out that the FDR bounds differ by a dependency factor which can be used as a correction quantity. We offer sparsity modifications and improved dependence tests which generalize the Benjamini-Yekutieli test and adaptive tests in the sense of Storey. As a byproduct, an early stopped test is presented in order to bound the number of rejections. The new procedures perform well for real genome data examples.

math.ST

Dynamic adaptive procedures that control the false discovery rate

In the multiple testing problem with independent tests, the classical linear step-up procedure controls the false discovery rate (FDR) at level $π_0α$, where $π_0$ is the proportion of true null hypotheses and $α$ is the target FDR level. Adaptive procedures can improve power by incorporating estimates of $π_0$, which typically rely on a tuning parameter. Fixed adaptive procedures set their tuning parameters before seeing the data and can be shown to control the FDR in finite samples. We develop theoretical results for dynamic adaptive procedures whose tuning parameters are determined by the data. We show that, if the tuning parameter is chosen according to a left-to-right stopping time rule, the corresponding dynamic adaptive procedure controls the FDR in finite samples. Examples include the recently proposed right-boundary procedure and the widely used lowest-slope procedure, among others. Simulation results show that the right-boundary procedure is more powerful than other dynamic adaptive procedures under independence and mild dependence conditions.

stat.ME

Detectability of nonparametric signals: higher criticism versus likelihood ratio

We study the signal detection problem in high dimensional noise data (possibly) containing rare and weak signals. Log-likelihood ratio (LLR) tests depend on unknown parameters, but they are needed to judge the quality of detection tests since they determine the detection regions. The popular Tukey's higher criticism (HC) test was shown to achieve the same completely detectable region as the LLR test does for different (mainly) parametric models. We present a novel technique to prove this result for very general signal models, including even nonparametric $p$-value models. Moreover, we address the following questions which are still pending since the initial paper of Donoho and Jin: What happens on the border of the completely detectable region, the so-called detection boundary? Does HC keep its optimality there? In particular, we give a complete answer for the heteroscedastic normal mixture model. As a byproduct, we give some new insights about the LLR test's behavior on the detection boundary by discussing, among others, Pitmans's asymptotic efficiency as an application of Le Cam's theory.

math.ST

On the consistency of adaptive multiple tests

Much effort has been done to control the "false discovery rate" (FDR) when $m$ hypotheses are tested simultaneously. The FDR is the expectation of the "false discovery proportion" $\text{FDP}=V/R$ given by the ratio of the number of false rejections $V$ and all rejections $R$. In this paper, we have a closer look at the FDP for adaptive linear step-up multiple tests. These tests extend the well known Benjamini and Hochberg test by estimating the unknown amount $m_0$ of the true null hypotheses. We give exact finite sample formulas for higher moments of the FDP and, in particular, for its variance. Using these allows us a precise discussion about the consistency of adaptive step-up tests. We present sufficient and necessary conditions for consistency on the estimators $\widehat m_0$ and the underlying probability regime. We apply our results to convex combinations of generalized Storey type estimators with various tuning parameters and (possibly) data-driven weights. The corresponding step-up tests allow a flexible adaptation. Moreover, these tests control the FDR at finite sample size. We compare these tests to the classical Benjamini and Hochberg test and discuss the advantages of it.

math.ST

Finite sample bounds for expected number of false rejections under martingale dependence with applications to FDR

Much effort has been made to improve the famous step up test of Benjamini and Hochberg given by linear critical values $\frac{iα}{n}$. It is pointed out by Gavrilov, Benjamini and Sarkar that step down multiple tests based on the critical values $β_i=\frac{iα}{n+1-i(1-α)}$ still control the false discovery rate (FDR) at the upper bound $α$ under basic independence assumptions. Since that result in not longer true for step up tests or dependent single tests, a big discussion about the corresponding FDR starts in the literature. The present paper establishes finite sample formulas and bounds for the FDR and the expected number of false rejections for multiple tests using critical values $β_i$ under martingale and reverse martingale dependence models. It is pointed out that martingale methods are natural tools for the treatment of local FDR estimators which are closely connected to the present coefficients $β_i.$ The martingale approach also yields new results and further inside for the special basic independence model.

math.ST

The False Discovery Rate (FDR) of Multiple Tests in a Class Room Lecture

Multiple tests are designed to test a whole collection of null hypotheses simultaneously. Their quality is often judged by the false discovery rate (FDR), i.e. the expectation of the quotient of the number of false rejections divided by the amount of all rejections. The widely cited Benjamini and Hochberg (BH) step up multiple test controls the FDR under various regularity assumptions. In this note we present a rapid approach to the BH step up and step down tests. Also sharp FDR inequalities are discussed for dependent p-values and examples and counter-examples are considered. In particular, the Bonferroni bound is sharp under dependence for control of the family-wise error rate.

math.ST

Inequalities for the false discovery rate (FDR) under dependence

Inequalities are key tools to prove FDR control of a multiple test. The present paper studies upper and lower bounds for the FDR under various dependence structures of p-values, namely independence, reverse martingale dependence and positive regression dependence on the subset (PRDS) of true null hypotheses. The inequalities are based on exact finite sample formulas which are also of interest for independent uniformly distributed p-values under the null. As applications the asymptotic worst case FDR of step up and step down tests coming from an non-decreasing rejection curve is established. In addition, new step up tests are established and necessary conditions for the FDR control are discussed. The reverse martingale models yield sharper FDR results than the PRDS models. Already in certain multivariate normal dependence models the familywise error rate of the Benjamini Hochberg step up test can be different from the desired level alpha. The second part of the paper is devoted to adaptive step up tests under dependence. The well-known Storey estimator is modified so that the corresponding step up test has finite sample control for various block wise dependent p-values. These results may be applied to dependent genome data. Within each chromosome the p-values may be reverse martingale dependent while the chromosomes are independent.

math.ST

Dynamic adaptive multiple tests with finite sample FDR control

The present paper introduces new adaptive multiple tests which rely on the estimation of the number of true null hypotheses and which control the false discovery rate (FDR) at level alpha for finite sample size. We derive exact formulas for the FDR for a large class of adaptive multiple tests which apply to a new class of testing procedures. In the following, generalized Storey estimators and weighted versions are introduced and it turns out that the corresponding adaptive step up and step down tests control the FDR. The present results also include particular dynamic adaptive step wise tests which use a data dependent weighting of the new generalized Storey estimators. In addition, a converse of the Benjamini Hochberg (1995) theorem is given. The Benjamini Hochberg (1995) test is the only "distribution free" step up test with FDR independent of the distribution of the p-values of false null hypotheses.

math.ST

Exponent dependence measures of survival functions and correlated frailty models

The present article studies survival analytic aspects of semiparametric copula dependence models with arbitrary univariate marginals. The underlying survival functions admit a representation via exponent measures which have an interpretation within the context of hazard functions. In particular, correlated frailty survival models are linked to copulas. Additionally, the relation to exponent measures of minumum-infinitely divisible distributions as well as to the Lévy measure of the Lévy-Khintchine formula is pointed out. The semiparametric character of the current analyses and the construction of survival times with dependencies of higher order are carried out in detail. Many examples including graphics give multifarious illustrations.

math.ST

Statistical likelihood methods in finance

It is known from previous work of the authors that non-negative arbitrage free price processes in finance can be described in terms of filtered likelihood processes of statistical experiments and vice versa. The present paper summarizes and outlines some similarities between finance and the statistical likelihood theory of Le Cam. Options are linked to statistical tests of the underlying experiments. In particular, some price formulas for options are expressed by the power of related tests. In special cases the dynamics of power functions for filtered likelihood processes can be used to establish trading strategies which lead to formulas for the Greeks Delta and Gamma. Moreover statistical arguments are then used to establish a discrete approximation of continuous time trading strategies. It is explained that Ito type financial models correspond to hazard based survival models in statistics. Also price processes given by a geometric fractional Brownian motion have a statistical counterpart in terms of the likelihood theory of Gaussian statistical experiments.

math.PR

The Convolution Theorem of Hajek and Le Cam - Revisited

The present paper establishes convolution theorems for regular estimators when the limit experiment is non-Gaussian or of infnite dimension with sparse parameter space. Applications are given for Gaussian shift experiments of infnite dimension, the Brownian motion signal plus noise model, Levy processes which are observed at discrete times and estimators of the endpoints of densities with jumps. The method of proof is also of interest for the classical convolution theorem of Hajek and Le Cam. As technical tool we present an elementary approach for the comparison of limit experiments on standard Borel spaces.

math.ST

Applications of the Likelihood Theory in Finance: Modelling and Pricing

This paper discusses the connection between mathematical finance and statistical modelling which turns out to be more than a formal mathematical correspondence. We like to figure out how common results and notions in statistics and their meaning can be translated to the world of mathematical finance and vice versa. A lot of similarities can be expressed in terms of LeCam's theory for statistical experiments which is the theory of the behaviour of likelihood processes. For positive prices the arbitrage free financial assets fit into filtered experiments. It is shown that they are given by filtered likelihood ratio processes. From the statistical point of view, martingale measures, completeness and pricing formulas are revisited. The pricing formulas for various options are connected with the power functions of tests. For instance the Black-Scholes price of a European option has an interpretation as Bayes risk of a Neyman Pearson test. Under contiguity the convergence of financial experiments and option prices are obtained. In particular, the approximation of Ito type price processes by discrete models and the convergence of associated option prices is studied. The result relies on the central limit theorem for statistical experiments, which is well known in statistics in connection with local asymptotic normal (LAN) families. As application certain continuous time option prices can be approximated by related discrete time pricing formulas.

math.ST