SearcharxivSearch

arXiv subjects

Bikram Karmakar

Publications and source records attributed to Bikram Karmakar.

9 recordsLinked to original sources

Causal mediation analysis for zero-inflated longitudinal data in the presence of treatment non-compliance and multiple mediators

Understanding whether a digital marketing campaign is effective is central to designing effective customer engagement strategies. We analyze a large-scale, longitudinal promotional email campaign conducted by a U.S.\ retailer to evaluate how value-added incentives, such as free shipping, compare with traditional price discounts in influencing customer purchasing behavior. The analysis is complicated by non-compliance, due to not opening emails, multiple longitudinal mediators, and zero-inflated mediators and purchase outcomes. To address these challenges, we develop a Bayesian causal mediation framework based on enriched Dirichlet process mixture models and estimate the causal estimands using a scalable G-computation algorithm. We show that analyses ignoring email-opening behavior substantially attenuate estimated effects. Value-added incentives consistently outperform price discounts, yielding higher estimated potential purchase amounts, with benefits accumulating over time. We design an individualized sequential emailing strategy that optimizes expected purchase count in the observed data.

stat.ME

Robustness and Efficiency of Rosenbaum's Rank-based Estimator in Randomized Trials: A Design-based Perspective

Mean-based estimators of causal effects in randomized experiments may behave poorly if the potential outcomes have a heavy tail or contain outliers. An alternative estimator proposed by Rosenbaum (1993) estimates a constant additive treatment effect by inverting a randomization test using ranks. We develop a design-based asymptotic theory for this rank-based estimator and study its robustness and efficiency properties. We show that Rosenbaum's estimator is robust against outliers with a breakdown point that uniformly dominates that of any weighted quantile estimator. When pretreatment covariates are available, a regression-adjusted version of Rosenbaum's estimator uses an agnostic linear regression on the covariates and bases inference on the ranks of residuals. Under mild integrability conditions, we show that this estimator is at most 13.6% less efficient, in the worst case, than the commonly used mean-based regression adjustment method proposed by Lin (2013); often outperforming it when the residuals have heavy tails. Moreover, under suitable assumptions, Rosenbaum's regression-adjusted estimator is at least as efficient as the unadjusted one. Finally, we initiate the study of Rosenbaum's estimator when the constant treatment effect assumption may be violated. To analyze the regression-adjusted estimator, we develop local asymptotics of rank statistics under the design-based framework, which may be of independent interest.

stat.ME

Using a Two-Parameter Sensitivity Analysis Framework to Efficiently Combine Randomized and Non-randomized Studies

Causal inference is vital for informed decision-making across fields such as biomedical research and social sciences. Randomized controlled trials (RCTs) are considered the gold standard for internal validity of inferences, whereas observational studies (OSs) often provide the opportunity for greater external validity. However, both data sources have inherent limitations preventing their use for broadly valid statistical inferences: RCTs may lack generalizability due to their selective eligibility criterion, and OSs are vulnerable to unobserved confounding. This paper proposes an innovative approach to integrate RCT and OS that borrows the other study's strengths to remedy each study's limitations. The method uses a novel triplet matching algorithm to align RCT and OS samples and a new two-parameter sensitivity analysis framework to quantify internal and external validity biases. This combined approach yields causal estimates that are more robust to hidden biases than OSs alone and provides reliable inferences about the treatment effect in the general population. We apply this method to investigate the effects of lactation on maternal health using a small RCT and a long-term observational health records dataset from the California National Primate Research Center. This application demonstrates the practical utility of our approach in generating scientifically sound and actionable causal estimates.

stat.ME

Degree of Interference: A General Framework For Causal Inference Under Interference

One core assumption typically adopted for valid causal inference is that of no interference between experimental units, i.e., the outcome of an experimental unit is unaffected by the treatments assigned to other experimental units. This assumption can be violated in real-life experiments, which significantly complicates the task of causal inference. As the number of potential outcomes increases, it becomes challenging to disentangle direct treatment effects from ``spillover'' effects. Current methodologies are lacking, as they cannot handle arbitrary, unknown interference structures to permit inference on causal estimands. We present a general framework to address the limitations of existing approaches. Our framework is based on the new concept of the ``degree of interference'' (DoI). The DoI is a unit-level latent variable that captures the latent structure of interference. We also develop a data augmentation algorithm that adopts a blocked Gibbs sampler and Bayesian nonparametric methodology to perform inferences on the estimands under our framework. We illustrate the DoI concept and properties of our Bayesian methodology via extensive simulation studies and an analysis of a randomized experiment investigating the impact of a cash transfer program for which interference is a critical concern. Ultimately, our framework enables us to infer causal effects without strong structural assumptions on interference.

stat.ME

Inferring the Effect of a Confounded Treatment by Calibrating Resistant Population's Variance

In a general set-up that allows unmeasured confounding, we show that the conditional average treatment effect on the treated can be identified as one of two possible values. Unlike existing causal inference methods, we do not require an exogenous source of variability in the treatment, e.g., an instrument or another outcome unaffected by the treatment. Instead, we require (a) a nondeterministic treatment assignment, (b) that conditional variances of the two potential outcomes are equal in the treatment group, and (c) a resistant population that was not exposed to the treatment or, if exposed, is unaffected by the treatment. Assumption (a) is commonly assumed in theoretical work, while (b) holds under fairly general outcome models. For (c), which is a new assumption, we show that a resistant population is often available in practice. We develop a large sample inference methodology and demonstrate our proposed method in a study of the effect of surface mining in central Appalachia on birth weight that finds a harmful effect.

stat.ME

Inference for a test-negative case-control study with added controls

Test-negative designs with added controls have recently been proposed to study COVID-19. An individual is test-positive or test-negative accordingly if they took a test for a disease but tested positive or tested negative. Adding a control group to a comparison of test-positives vs test-negatives is useful since additional comparison of test-positives vs controls can have potential biases different from the first comparison. Bonferroni correction ensures necessary type-I error control for these two comparisons done simultaneously. We propose two new methods for inference which have better interpretability and higher statistical power for these designs. These methods add a third comparison that is essentially independent of the first comparison, but our proposed second method often pays much less for these three comparisons than what a Bonferroni correction would pay for the two comparisons.

stat.ME

Statistical Validity and Consistency of Big Data Analytics: A General Framework

Informatics and technological advancements have triggered generation of huge volume of data with varied complexity in its management and analysis. Big Data analytics is the practice of revealing hidden aspects of such data and making inferences from it. Although storage, retrieval and management of Big Data seem possible through efficient algorithm and system development, concern about statistical consistency remains to be addressed in view of its specific characteristics. Since Big Data does not conform to standard analytics, we need proper modification of the existing statistical theory and tools. Here we propose, with illustrations, a general statistical framework and an algorithmic principle for Big Data analytics that ensure statistical accuracy of the conclusions. The proposed framework has the potential to push forward advancement of Big Data analytics in the right direction. The partition-repetition approach proposed here is broad enough to encompass all practical data analytic problems.

cs.DB

Univariate and data-depth based multivariate control charts using trimmed mean and winsorized standard deviation

Over the years, the most popularly used control chart for statistical process control has been Shewhart's $\bar{X}-S$ or $\bar{X}-R$ chart along with its multivariate generalizations. But, such control charts suffer from the lack of robustness. In this paper, we propose a modified and improved version of Shewhart chart, based on trimmed mean and winsorized variance that proves robust and more efficient. We have generalized this approach of ours with suitable modifications using depth functions for Multivariate control charts and EWMA charts as well. We have discussed the theoretical properties of our proposed statistics and have shown the efficiency of our methodology on univariate and multivariate simulated datasets. We have also compared our approach to the other popular alternatives to Shewhart Chart already proposed and established the efficacy of our methodology.

stat.CO

Test for the statistical significance of a treatment effect in the presence of hidden sub-populations

For testing the statistical significance of a treatment effect, we usually compare between two parts of a population, one is exposed to the treatment, and the other is not exposed to it. Standard parametric and nonparametric two-sample tests are often used for this comparison. But direct applications of these tests can yield misleading results, especially when the population has some hidden sub-populations, and the impact of this sub-population difference on the study variables dominates the treatment effect. This problem becomes more evident if these subpopulations have widely different proportions of representatives in the samples taken from these two parts, which are often referred to as the treatment group and the control group. In this article, we make an attempt to overcome this problem. Our propose methods use suitable clustering algorithms to find the hidden sub-populations and then eliminate the sub-population effect by using suitable transformations. Standard two-sample tests, when they are applied on the transformed data, yield better results. Some simulated and real data sets are analyzed to show the utility of the proposed methods.

stat.CO