Searcharxiv⌕ Search

arXiv subjects

Sanat K. Sarkar

Publications and source records attributed to Sanat K. Sarkar.

At least 19 recordsLinked to original sources

Further Results on Controlling the False Discovery Rate in Two-Sided Gaussian Mean Testing

The recent work of Sarkar and Zhang (2025) introduced Positive Tail Dependence Under the Null (PTDN) and developed Generalized Shifted Benjamini-Hochberg (BH) procedures for two-sided Gaussian $z$- and $t$-testing under known covariance structures. This paper develops further consequences of that framework. First, we derive explicit dependence-adaptive lower and upper bounds for the FDR of the original BH procedure in terms of the conditional variance parameters $τ_i=1-R_i^2$, where $R_i^2$ is the squared multiple correlation between the $i$th statistic and the remaining coordinates. These bounds recover the exact BH FDR under independence and provide finite-sample, covariance-specific information complementary to generic bounds. We also identify conditions under which the coordinate-specific calibration of shifted BH can provide a rejection advantage over the original BH procedure. Second, we consider the practically important setting in which the covariance matrix is unknown but an independent Wishart estimator is available. Using simultaneous lower confidence bounds for the $τ_i$'s, we construct a confidence-bound shifted BH procedure and establish finite-sample FDR control. To our knowledge, this is the first shifted-BH-type procedure with a finite-sample guarantee for two-sided Gaussian mean testing under a completely unknown covariance matrix estimated independently. Numerical studies illustrate the behavior of the covariance-adaptive bounds, the potential advantage of shifted BH over BH, and the performance of confidence-bound shifting under unknown covariance.

stat.ME↗

Dependence-Aware False Discovery Rate Control in Two-Sided Gaussian Mean Testing

This paper develops a general framework for controlling the false discovery rate (FDR) in multiple testing of Gaussian means against two-sided alternatives. The widely used Benjamini-Hochberg (BH) procedure provides exact FDR control under independence or conservative control under specific one-sided dependence structures, but its validity for correlated two-sided tests has remained an open question. We introduce the notion of positive left-tail dependence under the null (PLTDN), extending classical dependence assumptions to two-sided settings, and show that it ensures valid FDR control for BH-type procedures. Building on this framework, we propose a family of generalized shifted BH (GSBH) methods that incorporate correlation information through simple p-value adjustments. Simulation results demonstrate reliable FDR control and improved power across a range of dependence structures, while an application to an HIV gene expression dataset illustrates the practical effectiveness of the proposed approach.

stat.ME↗

On Controlling the False Discovery Rate in Multiple Testing of the Means of Correlated Normals Against Two-Sided Alternatives

This paper revisits the following open question in simultaneous testing of multivariate normal means against two-sided alternatives: Can the method of Benjamini and Hochberg (BH, 1995) control the false discovery rate (FDR) without imposing any dependence structure on the correlations? The answer to this question is generally believed to be yes, and is conjectured so in the literature since results of numerical studies investigating the question and reported in numerous papers strongly support it. No theoretical justification of this answer has yet been put forward in the literature, as far as we know. In this paper, we offer a partial proof of this conjecture. More specifically, we consider the following two settings - (i) the covariance matrix is known and (ii) the covariance matrix is an unknown scalar multiple of a known matrix - and prove that in each of these settings a BH-type stepup method based on some weighted versions of the original z- or t-test statistics controls the FDR.

math.ST↗

Local False Discovery Rate Based Methods for Multiple Testing of One-Way Classified Hypotheses

This paper continues the line of research initiated in Liu et. al. (2016) on developing a novel framework for multiple testing of hypotheses grouped in a one-way classified form using hypothesis-specific local false discovery rates (Lfdr's). It is built on an extension of the standard two-class mixture model from single to multiple groups, defining hypothesis-specific Lfdr as a function of the conditional Lfdr for the hypothesis given that it is within an important group and the Lfdr for the group itself and involving a new parameter that measures grouping effect. This definition captures the underlying group structure for the hypotheses belonging to a group more effectively than the standard two-class mixture model. Two new Lfdr based methods, possessing meaningful optimalities, are produced in their oracle forms. One, designed to control false discoveries across the entire collection of hypotheses, is proposed as a powerful alternative to simply pooling all the hypotheses into a single group and using commonly used Lfdr based method under the standard single-group two-class mixture model. The other is proposed as an Lfdr analog of the method of Benjamini and Bogomolov (2014) for selective inference. It controls Lfdr based measure of false discoveries associated with selecting groups concurrently with controlling the average of within-group false discovery proportions across the selected groups. Simulation studies and real-data application show that our proposed methods are often more powerful than their relevant competitors.

stat.ME↗

Adjusting the Benjamini-Hochberg method for controlling the false discovery rate in knockoff assisted variable selection

The knockoff-based multiple testing setup of Barber & Candes (2015) for variable selection in multiple regression where sample size is as large as the number of explanatory variables is considered. The method of Benjamini & Hochberg (1995) based on ordinary least squares estimates of the regression coefficients is adjusted to the setup, transforming it to a valid p-value based false discovery rate controlling method not relying on any specific correlation structure of the explanatory variables. Simulations and real data applications show that our proposed method that is agnostic to π0, the proportion of unimportant explanatory variables, and a data-adaptive version of it that uses an estimate of π0 are powerful competitors of the false discovery rate controlling method in Barber & Candes (2015).

stat.ME↗

Controlling the False Discovery Rate in Complex Multi-Way Classified Hypotheses

In this article, we propose a generalized weighted version of the well-known Benjamini-Hochberg (BH) procedure. The rigorous weighting scheme used by our method enables it to encode structural information from simultaneous multi-way classification as well as hierarchical partitioning of hypotheses into groups, with provisions to accommodate overlapping groups. The method is proven to control the False Discovery Rate (FDR) when the p-values involved are Positively Regression Dependent on the Subset (PRDS) of null p-values. A data-adaptive version of the method is proposed. Simulations show that our proposed methods control FDR at desired level and are more powerful than existing comparable multiple testing procedures, when the p-values are independent or satisfy certain dependence conditions. We apply this data-adaptive method to analyze a neuro-imaging dataset and understand the impact of alcoholism on human brain. Neuro-imaging data typically have complex classification structure, which have not been fully utilized in subsequent inference by previously proposed multiple testing procedures. With a flexible weighting scheme, our method is poised to extract more information from the data and use it to perform a more informed and efficient test of the hypotheses.

stat.ME↗

A grouped, selectively weighted false discovery rate procedure

False discovery rate (FDR) control in structured hypotheses testing is an important topic in simultaneous inference. Most existing methods that aim to utilize group structure among hypotheses either employ the groupwise mixture model or weight all p-values or hypotheses. Thus, their powers can be improved when the groupwise mixture model is inappropriate or when most groups contain only true null hypotheses. Motivated by this, we propose a grouped, selectively weighted FDR procedure, which we refer to as "sGBH". Specifically, without employing the groupwise mixture model, sGBH identifies groups of hypotheses of interest, weights p-values in each such group only, and tests only the selected hypotheses using the weighted p-values. The sGBH subsumes a standard grouped, weighted FDR procedure which we refer to as "GBH". We provide simple conditions to ensure the conservativeness of sGBH, together with empirical evidence on its much improved power over GBH. The new procedure is applied to a gene expression study.

stat.ME↗

A weighted FDR procedure under discrete and heterogeneous null distributions

Multiple testing with false discovery rate (FDR) control has been widely conducted in the ``discrete paradigm" where p-values have discrete and heterogeneous null distributions. However, in this scenario existing FDR procedures often lose some power and may yield unreliable inference, and for this scenario there does not seem to be an FDR procedure that partitions hypotheses into groups, employs data-adaptive weights and is non-asymptotically conservative. We propose a weighted FDR procedure for multiple testing in the discrete paradigm that efficiently adapts to both the heterogeneity and discreteness of p-value distributions. We theoretically justify the non-asymptotic conservativeness of the weighted FDR procedure under independence, and show via simulation studies that, for multiple testing based on p-values of Binomial test or Fisher's exact test, it is more powerful than six other procedures. The weighted FDR procedure is applied to a drug safety study and a differential methylation study based on discrete data, where it makes more discoveries than two existing methods.

stat.ME↗

On Benjamini-Hochberg procedure applied to mid p-values

Multiple testing with discrete p-values routinely arises in various scientific endeavors. However, procedures, including the false discovery rate (FDR) controlling Benjamini-Hochberg (BH) procedure, often used in such settings, being developed originally for p-values with continuous distributions, are too conservative, and so may not be as powerful as one would hope for. Therefore, improving the BH procedure by suitably adapting it to discrete p-values without losing its FDR control is currently an important path of research. This paper studies the FDR control of the BH procedure when it is applied to mid p-values and derive conditions under which it is conservative. Our simulation study reveals that the BH procedure applied to mid p-values may be conservative under much more general settings than characterized in this work, and that an adaptive version of the BH procedure applied to mid p-values is as powerful as an existing adaptive procedure based on randomized p-values.

stat.ME↗

Adapting BH to One- and Two-Way Classified Structures of Hypotheses

Multiple testing literature contains ample research on controlling false discoveries for hypotheses classified according to one criterion, which we refer to as one-way classified hypotheses. Although simultaneous classification of hypotheses according to two different criteria, resulting in two-way classified hypotheses, do often occur in scientific studies, no such research has taken place yet, as far as we know, under this structure. This article produces procedures, both in their oracle and data-adaptive forms, for controlling the overall false discovery rate (FDR) across all hypotheses effectively capturing the underlying one- or two-way classification structure. They have been obtained by using results associated with weighted Benjamini-Hochberg (BH) procedure in their more general forms providing guidance on how to adapt the original BH procedure to the underlying one- or two-way classification structure through an appropriate choice of the weights. The FDR is maintained non-asymptotically by our proposed procedures in their oracle forms under positive regression dependence on subset of null $p$-values (PRDS) and in their data-adaptive forms under independence of the $p$-values. Possible control of FDR for our data-adaptive procedures in certain scenarios involving dependent $p$-values have been investigated through simulations. The fact that our suggested procedures can be superior to contemporary practices has been demonstrated through their applications in simulated scenarios and to real-life data sets. While the procedures proposed here for two-way classified hypotheses are new, the data-adaptive procedure obtained for one-way classified hypotheses is alternative to and often more powerful than those proposed in Hu et al. (2010).

stat.ME↗

The Control of the False Discovery Rate in Fixed Sequence Multiple Testing

Controlling the false discovery rate (FDR) is a powerful approach to multiple testing. In many applications, the tested hypotheses have an inherent hierarchical structure. In this paper, we focus on the fixed sequence structure where the testing order of the hypotheses has been strictly specified in advance. We are motivated to study such a structure, since it is the most basic of hierarchical structures, yet it is often seen in real applications such as statistical process control and streaming data analysis. We first consider a conventional fixed sequence method that stops testing once an acceptance occurs, and develop such a method controlling the FDR under both arbitrary and negative dependencies. The method under arbitrary dependency is shown to be unimprovable without losing control of the FDR and unlike existing FDR methods; it cannot be improved even by restricting to the usual positive regression dependence on subset (PRDS) condition. To account for any potential mistakes in the ordering of the tests, we extend the conventional fixed sequence method to one that allows more but a given number of acceptances. Simulation studies show that the proposed procedures can be powerful alternatives to existing FDR controlling procedures. The proposed procedures are illustrated through a real data set from a microarray experiment.

stat.ME↗

Further results on controlling the false discovery proportion

The probability of false discovery proportion (FDP) exceeding $γ\in[0,1)$, defined as $γ$-FDP, has received much attention as a measure of false discoveries in multiple testing. Although this measure has received acceptance due to its relevance under dependency, not much progress has been made yet advancing its theory under such dependency in a nonasymptotic setting, which motivates our research in this article. We provide a larger class of procedures containing the stepup analog of, and hence more powerful than, the stepdown procedure in Lehmann and Romano [Ann. Statist. 33 (2005) 1138-1154] controlling the $γ$-FDP under similar positive dependence condition assumed in that paper. We offer better alternatives of the stepdown and stepup procedures in Romano and Shaikh [IMS Lecture Notes Monogr. Ser. 49 (2006a) 33-50, Ann. Statist. 34 (2006b) 1850-1873] using pairwise joint distributions of the null $p$-values. We generalize the notion of $γ$-FDP making it appropriate in situations where one is willing to tolerate a few false rejections or, due to high dependency, some false rejections are inevitable, and provide methods that control this generalized $γ$-FDP in two different scenarios: (i) only the marginal $p$-values are available and (ii) the marginal $p$-values as well as the common pairwise joint distributions of the null $p$-values are available, and assuming both positive dependence and arbitrary dependence conditions on the $p$-values in each scenario. Our theoretical findings are being supported through numerical studies.

math.ST↗

Applying multiple testing procedures to detect change in East African vegetation

The study of vegetation fluctuations gives valuable information toward effective land use and development. We consider this problem for the East African region based on the Normalized Difference Vegetation Index (NDVI) series from satellite remote sensing data collected between 1982 and 2006 over 8-kilometer grid points. We detect areas with significant increasing or decreasing monotonic vegetation changes using a multiple testing procedure controlling the mixed directional false discovery rate (mdFDR). Specifically, we use a three-stage directional Benjamini--Hochberg (BH) procedure with proven mdFDR control under independence and a suitable adaptive version of it. The performance of these procedures is studied through simulations before applying them to the vegetation data. Our analysis shows increasing vegetation in the Northern hemisphere as well as coastal Tanzania and generally decreasing Southern hemisphere vegetation trends, which are consistent with historical evidence.

stat.AP↗

Capturing the Severity of Type II Errors in High-Dimensional Multiple Testing

The severity of type II errors is frequently ignored when deriving a multiple testing procedure, even though utilizing it properly can greatly help in making correct decisions. This paper puts forward a theory behind developing a multiple testing procedure that can incorporate the type II error severity and is optimal in the sense of minimizing a measure of false non-discoveries among all procedures controlling a measure of false discoveries. The theory is developed under a general model allowing arbitrary dependence by taking a compound decision theoretic approach to multiple testing with a loss function incorporating the type II error severity. We present this optimal procedure in its oracle form and offer numerical evidence of its superior performance over relevant competitors.

stat.ME↗

On a generalized false discovery rate

The concept of $k$-FWER has received much attention lately as an appropriate error rate for multiple testing when one seeks to control at least $k$ false rejections, for some fixed $k\ge 1$. A less conservative notion, the $k$-FDR, has been introduced very recently by Sarkar [Ann. Statist. 34 (2006) 394--415], generalizing the false discovery rate of Benjamini and Hochberg [J. Roy. Statist. Soc. Ser. B 57 (1995) 289--300]. In this article, we bring newer insight to the $k$-FDR considering a mixture model involving independent $p$-values before motivating the developments of some new procedures that control it. We prove the $k$-FDR control of the proposed methods under a slightly weaker condition than in the mixture model. We provide numerical evidence of the proposed methods' superior power performance over some $k$-FWER and $k$-FDR methods. Finally, we apply our methods to a real data set.

math.ST↗

An adaptive step-down procedure with proven FDR control under independence

In this work we study an adaptive step-down procedure for testing $m$ hypotheses. It stems from the repeated use of the false discovery rate controlling the linear step-up procedure (sometimes called BH), and makes use of the critical constants $iq/[(m+1-i(1-q)]$, $i=1,...,m$. Motivated by its success as a model selection procedure, as well as by its asymptotic optimality, we are interested in its false discovery rate (FDR) controlling properties for a finite number of hypotheses. We prove this step-down procedure controls the FDR at level $q$ for independent test statistics. We then numerically compare it with two other procedures with proven FDR control under independence, both in terms of power under independence and FDR control under positive dependence.

math.ST↗

On the Simes inequality and its generalization

The Simes inequality has received considerable attention recently because of its close connection to some important multiple hypothesis testing procedures. We revisit in this article an old result on this inequality to clarify and strengthen it and a recently proposed generalization of it to offer an alternative simpler proof.

math.ST↗

Stepup procedures controlling generalized FWER and generalized FDR

In many applications of multiple hypothesis testing where more than one false rejection can be tolerated, procedures controlling error rates measuring at least $k$ false rejections, instead of at least one, for some fixed $k\ge 1$ can potentially increase the ability of a procedure to detect false null hypotheses. The $k$-FWER, a generalized version of the usual familywise error rate (FWER), is such an error rate that has recently been introduced in the literature and procedures controlling it have been proposed. A further generalization of a result on the $k$-FWER is provided in this article. In addition, an alternative and less conservative notion of error rate, the $k$-FDR, is introduced in the same spirit as the $k$-FWER by generalizing the usual false discovery rate (FDR). A $k$-FWER procedure is constructed given any set of increasing constants by utilizing the $k$th order joint null distributions of the $p$-values without assuming any specific form of dependence among all the $p$-values. Procedures controlling the $k$-FDR are also developed by using the $k$th order joint null distributions of the $p$-values, first assuming that the sets of null and nonnull $p$-values are mutually independent or they are jointly positively dependent in the sense of being multivariate totally positive of order two (MTP$_2$) and then discarding that assumption about the overall dependence among the $p$-values.

math.ST↗