SearcharxivSearch

arXiv subjects

Debashis Ghosh

Publications and source records attributed to Debashis Ghosh.

29 records · Page 2Linked to original sources

Analysis of regression discontinuity designs using censored data

In medical settings, treatment assignment may be determined by a clinically important covariate that predicts patients' risk of event. There is a class of methods from the social science literature known as regression discontinuity (RD) designs that can be used to estimate the treatment effect in this situation. Under certain assumptions, such an estimand enjoys a causal interpretation. However, few authors have discussed the use of RD for censored data. In this paper, we show how to estimate causal effects under the regression discontinuity design for censored data. The proposed estimation procedure employs a class of censoring unbiased transformations that includes inverse probability censored weighting and doubly robust transformation schemes. Simulation studies demonstrate the utility of the proposed methodology.

stat.ME

A Gaussian process framework for overlap and causal effect estimation with high-dimensional covariates

A powerful tool for the analysis of nonrandomized observational studies has been the potential outcomes model. Utilization of this framework allows analysts to estimate average treatment effects. This article considers the situation in which high-dimensional covariates are present and revisits the standard assumptions made in causal inference. We show that by employing a flexible Gaussian process framework, the assumption of strict overlap leads to very restrictive assumptions about the distribution of covariates, results for which can be characterized using classical results from Gaussian random measures as well as reproducing kernel Hilbert space theory. In addition, we propose a strategy for data-adaptive causal effect estimation that does not rely on the strict overlap assumption. These findings reveal the stringency that accompanies the use of the treatment positivity assumption in high-dimensional settings.

math.ST

A Simulation Based Dynamic Evaluation Framework for System-wide Algorithmic Fairness

We propose the use of Agent Based Models (ABMs) inside a reinforcement learning framework in order to better understand the relationship between automated decision making tools, fairness-inspired statistical constraints, and the social phenomena giving rise to discrimination towards sensitive groups. There have been many instances of discrimination occurring due to the applications of algorithmic tools by public and private institutions. Until recently, these practices have mostly gone unchecked. Given the large-scale transformation these new technologies elicit, a joint effort of social sciences and machine learning researchers is necessary. Much of the research has been done on determining statistical properties of such algorithms and the data they are trained on. We aim to complement that approach by studying the social dynamics in which these algorithms are implemented. We show how bias can be accumulated and reinforced through automated decision making, and the possibility of finding a fairness inducing policy. We focus on the case of recidivism risk assessment by considering simplified models of arrest. We find that if we limit our attention to what is observed and manipulated by these algorithmic tools, we may determine some blatantly unfair practices as fair, illustrating the advantage of analyzing the otherwise elusive property with a system-wide model. We expect the introduction of agent based simulation techniques will strengthen collaboration with social scientists, arriving at a better understanding of the social systems affected by technology and to hopefully lead to concrete policy proposals that can be presented to policymakers for a true systemic transformation.

cs.CY

Accuracy of the Epic Sepsis Prediction Model in a Regional Health System

Interest in an electronic health record-based computational model that can accurately predict a patient's risk of sepsis at a given point in time has grown rapidly in the last several years. Like other EHR vendors, the Epic Systems Corporation has developed a proprietary sepsis prediction model (ESPM). Epic developed the model using data from three health systems and penalized logistic regression. Demographic, comorbidity, vital sign, laboratory, medication, and procedural variables contribute to the model. The objective of this project was to compare the predictive performance of the ESPM with a regional health system's current Early Warning Score-based sepsis detection program.

stat.AP

Predictive Directions for Individualized Treatment Selection in Clinical Trials

In many clinical trials, individuals in different subgroups have experience differential treatment effects. This leads to individualized differences in treatment benefit. In this article, we introduce the general concept of predictive directions, which are risk scores motivated by potential outcomes considerations. These techniques borrow heavily from sufficient dimension reduction (SDR) and causal inference methodology. Under some conditions, one can use existing methods from the SDR literature to estimate the directions assuming an idealized complete data structure, which subsequently yields an obvious extension to clinical trial datasets. In addition, we generalize the direction idea to a nonlinear setting that exploits support vector machines. The methodology is illustrated with application to a series of colorectal cancer clinical trials.

stat.ME

Relaxed covariate overlap and margin-based causal effect estimation

In most nonrandomized observational studies, differences between treatment groups may arise not only due to the treatment but also because of the effect of confounders. Therefore, causal inference regarding the treatment effect is not as straightforward as in a randomized trial. To adjust for confounding due to measured covariates, a variety of methods based on the potential outcomes framework are used to estimate average treatment effects. One of the key assumptions is treatment positivity, which states that the probability of treatment is bounded away from zero and one for any possible combination of the confounders. Methods for performing causal inference when this assumption is violated are relatively limited. In this article, we discuss a new balance-related condition involving the convex hulls of treatment groups, which I term relaxed covariate overlap. An advantage of this concept is that it can be linked to a concept from machine learning, termed the margin. Introduction of relaxed covariate overlap leads to an approach in which one can perform causal inference in a three-step manner. The methodology is illustrated with two examples.

stat.ME

Optimal Kernel Combination for Test of Independence against Local Alternatives

Testing the independence between two random variables $x$ and $y$ is an important problem in statistics and machine learning, where the kernel-based tests of independence is focused to address the study of dependence recently. The advantage of the kernel framework rests on its flexibility in choice of kernel. The Hilbert-Schmidt Independence Criterion (HSIC) was shown to be equivalent to a class of tests, where the tests are based on different distance-induced kernel pairs. In this work, we propose to select the optimal kernel pair by considering local alternatives, and evaluate the efficiency using the quadratic time estimator of HSIC. The local alternative offers the advantage that the measure of efficiency do not depend on a particular alternative, and only requires the knowledge of the asymptotic null distribution of the test. We show in our experiments that the proposed strategy results in higher power than other existing kernel selection approaches.

stat.ME

Testing the disjunction hypothesis using Voronoi diagrams with applications to genetics

Testing of the disjunction hypothesis is appropriate when each gene or location studied is associated with multiple $p$-values, each of which is of individual interest. This can occur when more than one aspect of an underlying process is measured. For example, cancer researchers may hope to detect genes that are both differentially expressed on a transcriptomic level and show evidence of copy number aberration. Currently used methods of $p$-value combination for this setting are overly conservative, resulting in very low power for detection. In this work, we introduce a method to test the disjunction hypothesis by using cumulative areas from the Voronoi diagram of two-dimensional vectors of $p$-values. Our method offers much improved power over existing methods, even in challenging situations, while maintaining appropriate error control. We apply the approach to data from two published studies: the first aims to detect periodic genes of the organism Schizosaccharomyces pombe, and the second aims to identify genes associated with prostate cancer.

stat.ME

Equivalence of Kernel Machine Regression and Kernel Distance Covariance for Multidimensional Trait Association Studies

Associating genetic markers with a multidimensional phenotype is an important yet challenging problem. In this work, we establish the equivalence between two popular methods: kernel-machine regression (KMR), and kernel distance covariance (KDC). KMR is a semiparametric regression frameworks that models the covariate effects parametrically, while the genetic markers are considered non-parametrically. KDC represents a class of methods that includes distance covariance (DC) and Hilbert-Schmidt Independence Criterion (HSIC), which are nonparametric tests of independence. We show the equivalence between the score test of KMR and the KDC statistic under certain conditions. This result leads to a novel generalization of the KDC test that incorporates the covariates. Our contributions are three-fold: (1) establishing the equivalence between KMR and KDC; (2) showing that the principles of kernel machine regression can be applied to the interpretation of KDC; (3) the development of a broader class of KDC statistics, that the members are the quantities of different kernels. We demonstrate the proposals using simulation studies. Data from the Alzheimer's Disease Neuroimaging Initiative (ADNI) is used to explore the association between the genetic variants on gene \emph{FLJ16124} and phenotypes represented in 3D structural brain MR images adjusting for age and gender. The results suggest that SNPs of \emph{FLJ16124} exhibit strong pairwise interaction effects that are correlated to the changes of brain region volumes.

stat.ML

Multiple Comparison Procedures for Neuroimaging Genomewide Association Studies

Recent research in neuroimaging has focused on assessing associations between genetic variants that are measured on a genomewide scale and brain imaging phenotypes. A large number of works in the area apply massively univariate analyses on a genomewide basis to find single nucleotide polymorphisms that influence brain structure. In this paper, we propose using various dimensionality reduction methods on both brain structural MRI scans and genomic data, motivated by the Alzheimer's Disease Neuroimaging Initiative (ADNI) study. We also consider a new multiple testing adjustment method and compare it with two existing false discovery rate (FDR) adjustment methods. The simulation results suggest an increase in power for the proposed method. The real data analysis suggests that the proposed procedure is able to find associations between genetic variants and brain volume differences that offer potentially new biological insights.

stat.CO

Multiple testing procedures under confounding

While multiple testing procedures have been the focus of much statistical research, an important facet of the problem is how to deal with possible confounding. Procedures have been developed by authors in genetics and statistics. In this chapter, we relate these proposals. We propose two new multiple testing approaches within this framework. The first combines sensitivity analysis methods with false discovery rate estimation procedures. The second involves construction of shrinkage estimators that utilize the mixture model for multiple testing. The procedures are illustrated with applications to a gene expression profiling experiment in prostate cancer.

stat.ME