SearcharxivSearch

arXiv subjects

Ritoban Kundu

Publications and source records attributed to Ritoban Kundu.

6 recordsLinked to original sources

Optimal Treatment Policy Estimation for Recurrent Events with a Competing Terminal Event: An Instrumented Difference-in-Differences Approach

Learning reproducible and generalizable optimal treatment policies for chronic diseases requires large, representative populations with long-term follow-up. Administrative health data provide a natural starting point, but their use is often limited by unmeasured confounding. We address this by proposing a novel framework based on Instrumented Difference-in-Differences (iDID) to estimate optimal policies for recurrent event outcomes subject to a terminating event. The iDID design is particularly useful in this setting because it leverages policy-induced treatment variation while allowing for persistent unmeasured differences across populations, relying on assumptions that are more plausible for administrative health data than those required by conventional IV or DID approaches. A key feature of our approach is that it explicitly addresses the fundamental challenge of avoiding policies that trivially reduce recurrent adverse events by increasing mortality. We derive two distinct Inverse Probability Weighted identifications and develop a multiply robust estimator that achieves consistency if any one of several subsets of nuisance models is correctly specified. We establish the estimator's consistency and asymptotic normality through large-sample theory and demonstrate its superior finite-sample performance over existing methods via simulation. Finally, we apply this framework to a national Medicare dataset to optimize first-line Type 2 Diabetes strategies, specifically targeting the minimization of disease-related hospitalizations while accounting for survival.

stat.ME

Identification, Estimation, and Inference for Sequential Causally Ordered Mediation Pathways

Mediation analysis plays an essential role in uncovering the mechanisms by which an exposure influences an outcome through intermediate pathways. While methodological advances for single-mediator settings are well established, rigorous tools for handling multiple, sequentially ordered mediators remain underdeveloped. Such settings are common in applications like longitudinal cohort studies, where exposures operate through complex chains of mediators over time. In this paper, we establish a general framework for sequentially ordered mediators that enables the identification and formal decomposition of the total effect into component path-specific effects. We also develop estimation procedures for mediation estimands with both continuous and categorical outcomes. Furthermore, we introduce a new testing strategy to conduct inference using a studentized statistic combined with data-splitting. This approach achieves valid Type I error control under the composite null across diverse data-generating mechanisms. Through extensive simulations and applications to two large-scale empirical studies, we demonstrate that the proposed methodology provides reliable estimation, valid inference, and improved power for discovering novel mediation pathways.

stat.ME

Network Structural Equation Models for Causal Mediation and Spillover Effects

Social network interference induces complex dependencies where a unit's outcome is influenced not only by its own exposure and mediator but also by those of connected neighbors. In such settings, a significant challenge lies in distinguishing direct exposure effects from interference-driven spillover effects, and further separating these from indirect effects mediated by intermediate variables. To address this, we propose a theoretical framework utilizing structural graphical models. Central to our approach is the Random Effects Network Structural Equation Model (REN-SEM), which extends the exposure mapping paradigm to capture these multifaceted spillover and mediation mechanisms while accounting for latent dependencies within mediators and outcomes. We establish general identification conditions and derive decomposition formulas for six distinct mechanistic estimands. Furthermore, for the class of Linear REN-SEMs, we develop a maximum likelihood estimation framework and establish a rigorous asymptotic theory tailored to non-i.i.d. network data, proving the consistency of our estimators and the validity of the variance estimates. The robustness and practical utility of our methodology are demonstrated through simulation experiments and an analysis of the Twitch Gamers Network, underscoring its effectiveness in quantifying intricate network-mediated exposure effects.

stat.ME

A Doubly Robust Framework for Addressing Outcome-Dependent Selection Bias in Multi-Cohort EHR Studies

Selection bias can hinder accurate estimation of association parameters in binary disease risk models using non-probability samples like electronic health records (EHRs). The issue is compounded when participants are recruited from multiple clinics/centers with varying selection mechanisms that may depend on the disease/outcome of interest. Traditional inverse-probability-weighted (IPW) methods, based on constructed parametric selection models, often struggle with misspecifications when selection mechanisms vary across cohorts. This paper introduces a new Joint Augmented Inverse Probability Weighted (JAIPW) method, which integrates individual-level data from multiple cohorts collected under potentially outcome-dependent selection mechanisms, with data from an external probability sample. JAIPW offers double robustness by incorporating a flexible auxiliary score model to address potential misspecifications in the selection models. We outline the asymptotic properties of the JAIPW estimator, and our simulations reveal that JAIPW achieves up to six times lower relative bias and five times lower root mean square error (RMSE) compared to the best performing joint IPW methods under scenarios with misspecified selection models. Applying JAIPW to the Michigan Genomics Initiative (MGI), a multi-clinic EHR-linked biobank, combined with external national probability samples, resulted in cancer-sex association estimates closely aligned with national benchmark estimates. We also analyzed the association between cancer and polygenic risk scores (PRS) in MGI to illustrate a situation where the exposure variable is not measured in the external probability sample.

stat.ME

Copula Structural Equation Models for Mediation Pathway Analysis

Structural equation models (SEMs) are fundamental to causal mediation pathway discovery. However, traditional SEM approaches often rely on \emph{ad hoc} model specifications when handling complex data structures such as mixed data types or non-normal data in which Gaussian assumptions for errors are rather restrictive. The invocation of copula dependence modeling methods to extend the classical linear SEMs mitigates several of key technical limitations, offering greater modeling flexibility to analyze non-Gaussian data. This paper presents a selective review of major developments in this area, highlighting recent advancements and their methodological implications.

stat.ME

A Framework for Understanding Selection Bias in Real-World Healthcare Data

Using administrative patient-care data such as Electronic Health Records (EHR) and medical/ pharmaceutical claims for population-based scientific research has become increasingly common. With vast sample sizes leading to very small standard errors, researchers need to pay more attention to potential biases in the estimates of association parameters of interest, specifically to biases that do not diminish with increasing sample size. Of these multiple sources of biases, in this paper, we focus on understanding selection bias. We present an analytic framework using directed acyclic graphs for guiding applied researchers to dissect how different sources of selection bias may affect estimates of the association between a binary outcome and an exposure (continuous or categorical) of interest. We consider four easy-to-implement weighting approaches to reduce selection bias with accompanying variance formulae. We demonstrate through a simulation study when they can rescue us in practice with analysis of real world data. We compare these methods using a data example where our goal is to estimate the well-known association of cancer and biological sex, using EHR from a longitudinal biorepository at the University of Michigan Healthcare system. We provide annotated R codes to implement these weighted methods with associated inference.

stat.ME