SearcharxivSearch

arXiv subjects

Shinjini Nandi

Publications and source records attributed to Shinjini Nandi.

4 recordsLinked to original sources

Leveraging the group structure of hypotheses for more powerful multiple testing with FDR control for the filtered rejection set

Modern biological studies often involve testing many hypotheses organized in a group or a hierarchical structure, such as a directed acyclic graph (DAG). In these studies, researchers often wish to control the false discovery rate (FDR) after filtering the discoveries to obtain interpretable results. For addressing this goal, Katsevich, Sabatti, and Bogomolov (2023, Journal of the American Statistical Association, 118(541), 165-176) developed a general method, Focused BH, that guarantees FDR control for the filtered rejection set for a pre-specified filter, under certain assumptions. We propose improving the power of Focused BH by adapting it to group or hierarchical structures of hypotheses using data-dependent weights. The general method incorporating such weights is referred to as Weighted Focused BH (WFBH). For DAG-structured hypotheses, we propose a variant of WFBH, which can gain power by being adaptive to the DAG structure, and by exploiting the logical relationships among the hypotheses. We prove that WFBH with weights that were proposed to adapt the Benjamini-Hochberg procedure to different group structures, as well as its proposed variant for testing DAG-structured hypotheses, control the post-filtering FDR under certain assumptions. Through simulations, we demonstrate that the latter variant is robust to deviations from these assumptions and can be considerably more powerful than comparable methods. Finally, we elucidate its practical use by applying it to real datasets from microbiome and gene expression studies.

stat.ME

Joint Fairness Model with Applications to Risk Predictions for Under-represented Populations

In data collection for predictive modeling, under-representation of certain groups, based on gender, race/ethnicity, or age, may yield less-accurate predictions for these groups. Recently, this issue of fairness in predictions has attracted significant attention, as data-driven models are increasingly utilized to perform crucial decision-making tasks. Existing methods to achieve fairness in the machine learning literature typically build a single prediction model in a manner that encourages fair prediction performance for all groups. These approaches have two major limitations: i) fairness is often achieved by compromising accuracy for some groups; ii) the underlying relationship between dependent and independent variables may not be the same across groups. We propose a Joint Fairness Model (JFM) approach for logistic regression models for binary outcomes that estimates group-specific classifiers using a joint modeling objective function that incorporates fairness criteria for prediction. We introduce an Accelerated Smoothing Proximal Gradient Algorithm to solve the convex objective function, and present the key asymptotic properties of the JFM estimates. Through simulations, we demonstrate the efficacy of the JFM in achieving good prediction performance and across-group parity, in comparison with the single fairness model, group-separate model, and group-ignorant model, especially when the minority group's sample size is small. Finally, we demonstrate the utility of the JFM method in a real-world example to obtain fair risk predictions for under-represented older patients diagnosed with coronavirus disease 2019 (COVID-19).

stat.AP

Controlling the False Discovery Rate in Complex Multi-Way Classified Hypotheses

In this article, we propose a generalized weighted version of the well-known Benjamini-Hochberg (BH) procedure. The rigorous weighting scheme used by our method enables it to encode structural information from simultaneous multi-way classification as well as hierarchical partitioning of hypotheses into groups, with provisions to accommodate overlapping groups. The method is proven to control the False Discovery Rate (FDR) when the p-values involved are Positively Regression Dependent on the Subset (PRDS) of null p-values. A data-adaptive version of the method is proposed. Simulations show that our proposed methods control FDR at desired level and are more powerful than existing comparable multiple testing procedures, when the p-values are independent or satisfy certain dependence conditions. We apply this data-adaptive method to analyze a neuro-imaging dataset and understand the impact of alcoholism on human brain. Neuro-imaging data typically have complex classification structure, which have not been fully utilized in subsequent inference by previously proposed multiple testing procedures. With a flexible weighting scheme, our method is poised to extract more information from the data and use it to perform a more informed and efficient test of the hypotheses.

stat.ME

Adapting BH to One- and Two-Way Classified Structures of Hypotheses

Multiple testing literature contains ample research on controlling false discoveries for hypotheses classified according to one criterion, which we refer to as one-way classified hypotheses. Although simultaneous classification of hypotheses according to two different criteria, resulting in two-way classified hypotheses, do often occur in scientific studies, no such research has taken place yet, as far as we know, under this structure. This article produces procedures, both in their oracle and data-adaptive forms, for controlling the overall false discovery rate (FDR) across all hypotheses effectively capturing the underlying one- or two-way classification structure. They have been obtained by using results associated with weighted Benjamini-Hochberg (BH) procedure in their more general forms providing guidance on how to adapt the original BH procedure to the underlying one- or two-way classification structure through an appropriate choice of the weights. The FDR is maintained non-asymptotically by our proposed procedures in their oracle forms under positive regression dependence on subset of null $p$-values (PRDS) and in their data-adaptive forms under independence of the $p$-values. Possible control of FDR for our data-adaptive procedures in certain scenarios involving dependent $p$-values have been investigated through simulations. The fact that our suggested procedures can be superior to contemporary practices has been demonstrated through their applications in simulated scenarios and to real-life data sets. While the procedures proposed here for two-way classified hypotheses are new, the data-adaptive procedure obtained for one-way classified hypotheses is alternative to and often more powerful than those proposed in Hu et al. (2010).

stat.ME