SearcharxivSearch

arXiv subjects

Zach Shahn

Publications and source records attributed to Zach Shahn.

At least 19 recordsLinked to original sources

As Good as it Gets: Bounds for Oracle Time-Varying Treatment Strategies

Much causal inference research is focused on methods for optimizing dynamic treatment regimes (Murphy, 2003; Robins, 2004; Schulte et al., 2015), which are rules for deciding which treatments should be assigned and when based on evolving history. There is a certain optimism underlying this endeavor that with enough tinkering we might realize consequential improvements. Another strand of research, previously confined to the point exposure setting, considers bounds on how well any individualized treatment rule could possibly do. Here, we extend to the time-varying setting sharp bounds on the performance of an oracle strategy that selects the best treatment regime for each subject based on their unobserved potential outcomes or `response type'. For binary outcomes, the lower bound (assuming higher is better) is simply the expected outcome attained by the optimal treatment regime based on observed history. For continuous outcomes, the lower bound may strictly exceed the maximal observed covariate based value. In the continuous setting, we also consider bounds on the CDF of oracle continuous potential outcomes.

math.ST

Trust Me, I'm a Doctor?

Clinical trials usually target average treatment effects, but treatment decisions are made for individuals. This tension motivates a common criticism of evidence-based medicine: a treatment that is beneficial on average may be inappropriate for a particular patient, and skilled physicians may outperform rigid adherence to the strategy that performed best in a randomized trial. We consider how randomized and observational data from the same target population can be used to assess that possibility. Specifically, we study settings in which trial and/or observational data yield estimates of average outcomes under `treat all', `treat none', and usual care strategies in the population of interest. We derive sharp bounds on the proportion of encounters with physicians whose personal strategies outperform the better of `treat all' or `treat none' under the assumption that no physician's strategy is worse than always choosing the worse performing of `treat all' and `treat none'. These results clarify when clinical data support relying on physician discretion over the trial-average recommendation and when stronger justification is required.

stat.AP

Structural Nested Mean Models for Modified Treatment Policies

There is a growing literature on estimating effects of treatment strategies based on the natural treatment that would have been received in the absence of intervention, often dubbed `modified treatment policies' (MTPs). MTPs are sometimes of interest because they are more realistic than interventions setting exposure to an ideal level for all members of a population. In the general time-varying setting, Richardson and Robins (2013) provided exchangeability conditions for nonparametric identification of MTP effects that could be deduced from Single World Intervention Graphs (SWIGs). Diaz (2023) provided multiply robust estimators under these identification assumptions that allow for machine learning nuisance regressions. In this paper, we fill a remaining gap by extending Structural Nested Mean Models (SNMMs) to MTP settings, which enables characterization of (time-varying) heterogeneity of MTP effects. We do this both under the exchangeability assumptions of Richardson and Robins (2013) and under parallel trends assumptions, which enables investigation of (time-varying heterogeneous) MTP effects in the presence of some unobserved confounding.

stat.ME

Identification and Estimation of Joint Potential Outcome Distributions from a Single Study

Most causal inference methods focus on estimating marginal average treatment effects, but many important causal estimands depend on the joint distribution of potential outcomes, including the probability of causation and proportions benefiting from or harmed by treatment. Wu et al (2025) recently established nonparametric identification of this joint distribution for categorical outcomes under binary treatment by leveraging variation across multiple studies. We demonstrate that their multi-study framework can be implemented within a single study by using a baseline covariate that is associated with untreated potential outcomes but does not modify treatment effects conditional on those outcomes. This reframing substantially broadens the practical applicability of their results, as it eliminates the need for multiple independent datasets and gives analysts control over covariate selection to satisfy key identifying assumptions. We provide complete identification and estimation theory for the single-study setting, including a Neyman-orthogonal estimator for cases where the conditional independence assumption only holds after adjusting for covariates. However, we argue that only in unusual settings would it be even theoretically possible for the identifying assumptions to hold exactly, making sensitivity analysis particularly important. We validate the estimator in a simulation and apply it to data from a large field experiment assessing the effect of mailings on voter turnout.

stat.ME

A Note on The Rationale Behind Using Parental Longevity as a Proxy in Mendelian Randomization Studies

In many cohorts (such as the UK Biobank) on which Mendelian Randomization studies are routinely performed, data on participants' longevity is inadequate as the majority of participants are still living. To nevertheless estimate effects on longevity, it is increasingly common for researchers to substitute participants' `parental attained age', i.e. parental lifespan or current age (which is routinely collected in UK Biobank), as a proxy outcome. The common approach to performing this clever trick appears to be based on a solid understanding of its underlying assumptions. However, we have not seen these assumptions (or the causal effects whose identification they enable) clearly stated anywhere in the literature. In this note, we fill that gap.

stat.ME

Using causal diagrams to assess parallel trends in difference-in-differences studies

Difference-in-differences (DID) is popular because it can allow for unmeasured confounding when the key assumption of parallel trends holds. However, there exists little guidance on how to decide a priori whether this assumption is reasonable. We attempt to develop such guidance by considering the relationship between a causal diagram and the parallel trends assumption. This is challenging because parallel trends is scale-dependent and causal diagrams are generally scale-independent. We develop conditions under which, given a nonparametric causal diagram, one can reject or fail to reject parallel trends. In particular, we adopt a linear faithfulness assumption, which states that all graphically connected variables are correlated, and which is often reasonable in practice. We show that parallel trends can be rejected if either (i) the treatment is affected by pre-treatment outcomes, or (ii) there exist unmeasured confounders for the effect of treatment on pre-treatment outcomes that are not confounders for the post-treatment outcome, or vice versa (more precisely, the two outcomes possess distinct minimally sufficient sets). We also argue that parallel trends should be strongly questioned if (iii) the pre-treatment outcomes affect the post-treatment outcomes (though the two can be correlated) since there exist reasonable semiparametric models in which such an effect violates parallel trends. When (i-iii) are absent, a necessary and sufficient condition for parallel trends is that the association between the common set of confounders and the potential outcomes is constant on an additive scale, pre- and post-treatment. These conditions are similar to, but more general than, those previously derived in linear structural equations models. We discuss our approach in the context of the effect of Medicaid expansion under the U.S. Affordable Care Act on health insurance coverage rates.

stat.ME

Generalizing Difference-in-Differences to Non-Canonical Settings: Identifying an Array of Estimands

Consider a general setting in which data on an outcome is collected in two `groups' at two time periods, with certain group-periods deemed `treated' and others `untreated'. A special case is the canonical Difference-in-Differences (DiD) setting in which one group is treated only in the second period while the other is treated in neither period. Then it is well known that under a parallel trends assumption across the two groups the classic DiD formula (subtracting the average change in outcome across periods in the treated group by the average change in the outcome across periods in the untreated group) identifies the average treatment effect on the treated in the second period. But other relations between group, period, and treatment are possible. For example, the groups might be demographic (or other baseline covariate) categories with all units in both groups treated in the second period and none treated in the first, i.e. a pre-post design. Or one group might be treated in both periods while the other is treated in neither. Furthermore, other parallel trends assumptions under other treatment regimes are possible. For example, we could assume the two groups' potential outcomes would evolve in parallel under a regime of `do not switch treatment in the second period'. In fact, there is a literal array of data structures and parallel trends assumptions. The difference between the changes in outcomes of the two groups, which we dub the `group DiD' (gDiD) formula, identifies different causal estimands depending on the data structure and parallel trends assumption adopted. Here, we determine under which combinations of data structure and assumptions the gDiD formula identifies meaningful causal estimands. We also explore when parallel trends assumptions are amenable to empirical check or structural justification via Single World Intervention Graphs.

stat.ME

Structural Nested Mean Models Under Parallel Trends with Interference

Despite the common occurrence of interference in Difference-in-Differences (DiD) applications, standard DiD methods rely on an assumption that interference is absent, and comparatively little work has considered how to accommodate and learn about spillover effects within a DiD framework. Here, we extend the `DiD-SNMMs' of Shahn et al (2022) to accommodate interference in a time-varying DiD setting. Doing so enables estimation of a richer set of effects than previous DiD approaches. For example, DiD-SNMMs do not assume the absence of spillover effects after direct exposures and can model how effects of direct or indirect (i.e. spillover) exposures depend on past and concurrent (direct or indirect) exposure and covariate history. We consider both cluster and network interference structures and illustrate the methodology in simulations and an application to effects of Medicaid expansion on uninsurance rates.

stat.ME

Estimating Heterogeneous Treatment Effects on Survival Outcomes Using Counterfactual Censoring Unbiased Transformations

Methods for estimating heterogeneous treatment effects (HTE) from observational data have largely focused on continuous or binary outcomes, with less attention paid to survival outcomes and almost none to settings with competing risks. In this work, we develop censoring unbiased transformations (CUTs) for survival outcomes both with and without competing risks. After converting time-to-event outcomes using these CUTs, direct application of HTE learners for continuous outcomes yields consistent estimates of heterogeneous cumulative incidence effects, total effects, and separable direct effects. Our CUTs enable application of a much larger set of state of the art HTE learners for censored outcomes than had previously been available, especially in competing risks settings. We provide generic model-free learner-specific oracle inequalities bounding the finite-sample excess risk. The oracle efficiency results depend on the oracle selector and estimated nuisance functions from all steps involved in the transformation. We demonstrate the empirical performance of the proposed methods in simulation studies.

stat.ME

Subgroup Difference in Differences to Identify Effect Modification Without a Control Group

Suppose it is of interest to characterize effect heterogeneity of an intervention across levels of a baseline covariate using only pre- and post- intervention outcome measurements from those who received the intervention, i.e. with no control group. For example, a researcher concerned with equity may wish to ascertain whether a minority group benefited less from an intervention than the majority group. We introduce the `subgroup parallel trends' assumption that the counterfactual untreated outcomes in each subgroup of interest follow parallel trends pre- and post- intervention. Under the subgroup parallel trends assumption, it is straightforward to show that a simple `subgroup difference in differences' (SDiD) expression (i.e., the average pre/post outcome difference in one subgroup subtracted by the average pre/post outcome difference in the other subgroup) identifies the difference between the intervention's effects in the two subgroups. This difference in effects across subgroups is identified even though the conditional effects in each subgroup are not. The subgroup parallel trends assumption is not stronger than the standard parallel trends assumption across treatment groups when a control group is available, and there are circumstances where it is more plausible. Thus, when effect modification by a baseline covariate is of interest, researchers might consider SDiD whether or not a control group is available.

stat.ME

Efficient estimation of weighted cumulative treatment effects by double/debiased machine learning

In empirical studies with time-to-event outcomes, investigators often leverage observational data to conduct causal inference on the effect of exposure when randomized controlled trial data is unavailable. Model misspecification and lack of overlap are common issues in observational studies, and they often lead to inconsistent and inefficient estimators of the average treatment effect. Estimators targeting overlap weighted effects have been proposed to address the challenge of poor overlap, and methods enabling flexible machine learning for nuisance models address model misspecification. However, the approaches that allow machine learning for nuisance models have not been extended to the setting of weighted average treatment effects for time-to-event outcomes when there is poor overlap. In this work, we propose a class of one-step cross-fitted double/debiased machine learning estimators for the weighted cumulative causal effect as a function of restriction time. We prove that the proposed estimators are consistent, asymptotically linear, and reach semiparametric efficiency bounds under regularity conditions. Our simulations show that the proposed estimators using nonparametric machine learning nuisance models perform as well as established methods that require correctly-specified parametric nuisance models, illustrating that our estimators mitigate the need for oracle parametric nuisance models. We apply the proposed methods to real-world observational data from a UK primary care database to compare the effects of anti-diabetic drugs on cancer clinical outcomes.

stat.ME

Bias Formulas for Violations of Proximal Identification Assumptions

Causal inference from observational data often rests on the unverifiable assumption of no unmeasured confounding. Recently, Tchetgen Tchetgen and colleagues have introduced proximal inference to leverage negative control outcomes and exposures as proxies to adjust for bias from unmeasured confounding. However, some of the key assumptions that proximal inference relies on are themselves empirically untestable. Additionally, the impact of violations of proximal inference assumptions on the bias of effect estimates is not well understood. In this paper, we derive bias formulas for proximal inference estimators under a linear structural equation model data generating process. These results are a first step toward sensitivity analysis and quantitative bias analysis of proximal inference estimators. While limited to a particular family of data generating processes, our results may offer some more general insight into the behavior of proximal inference estimators.

math.ST

When Do Outcome Driven Treatments Break Parallel Trends?

Under what circumstances is it a threat to the parallel trends assumption required for Difference in Differences (DiD) studies if treatment decisions are based on past values of the outcome? We explore via simulation studies whether parallel trends holds across a grid of data generating processes generally conducive to parallel trends (random walk, Hidden Markov Model, and constant direct additive confounding), study designs (never treated, not yet treated, or later treated control groups), and outcome responsiveness of treatment (yes or no). We interpret the upshot of our simulation results to be that parallel trends is typically not a credible assumption when treatments are influenced by past outcomes. This is due to a combination of regression to the mean and selection on future treatment values, depending on the control group. Since timing of treatment initiation is frequently influenced by past outcomes when the treatment is targeted at the outcome, perhaps DiD is generally better suited for studying unintended consequences of interventions?

stat.ME

Structural Nested Mean Models Under Parallel Trends Assumptions

We link and extend two approaches to estimating time-varying treatment effects on repeated continuous outcomes--time-varying Difference in Differences (DiD; see Roth et al. (2023) and Chaisemartin et al. (2023) for reviews) and Structural Nested Mean Models (SNMMs; see Vansteelandt and Joffe (2014) for a review). In particular, we show that SNMMs, previously known to be nonparametrically identified under a no unobserved confounding assumption, are also identified under a conditional parallel trends assumption similar to those typically used to justify time-varying DiD methods (but more amenable to time-varying confounding). Because SNMMs model a broader set of causal estimands, our results allow practitioners of time-varying DiD approaches to address additional types of substantive questions under similar assumptions. SNMMs enable estimation of time-varying effect heterogeneity, lasting effects of a `blip' of treatment at a single time point, effects of sustained interventions (possibly on continuous or multi-dimensional treatments) when treatment repeatedly changes value in the data, controlled direct effects, effects of dynamic treatment strategies that depend on covariate history, and more. We provide a method for sensitivity analysis to violations of our parallel trends assumption. We further explain how to estimate optimal treatment regimes via optimal regime SNMMs under parallel trends assumptions plus an assumption that there is no effect modification by unobserved confounders. Finally, we illustrate our methods with real data applications estimating effects of Medicaid expansion on uninsurance rates, effects of floods on flood insurance take-up, and effects of sustained changes in temperature on crop yields.

stat.ME

Blending Knowledge in Deep Recurrent Networks for Adverse Event Prediction at Hospital Discharge

Deep learning architectures have an extremely high-capacity for modeling complex data in a wide variety of domains. However, these architectures have been limited in their ability to support complex prediction problems using insurance claims data, such as readmission at 30 days, mainly due to data sparsity issue. Consequently, classical machine learning methods, especially those that embed domain knowledge in handcrafted features, are often on par with, and sometimes outperform, deep learning approaches. In this paper, we illustrate how the potential of deep learning can be achieved by blending domain knowledge within deep learning architectures to predict adverse events at hospital discharge, including readmissions. More specifically, we introduce a learning architecture that fuses a representation of patient data computed by a self-attention based recurrent neural network, with clinically relevant features. We conduct extensive experiments on a large claims dataset and show that the blended method outperforms the standard machine learning approaches.

cs.LG

A Formal Causal Interpretation of the Case-Crossover Design

The case-crossover design (Maclure, 1991) is widely used in epidemiology and other fields to study causal effects of transient treatments on acute outcomes. However, its validity and causal interpretation have only been justified under informal conditions. Here, we place the design in a formal counterfactual framework for the first time. Doing so helps to clarify its assumptions and interpretation. In particular, when the treatment effect is non-null, we identify a previously unnoticed bias arising from common causes of the outcome at different person-times. We analytically characterize the direction and size of this bias and demonstrate its potential importance with a simulation. We also use our derivation of the limit of the case-crossover estimator to analyze its sensitivity to treatment effect heterogeneity, a violation of one of the informal criteria for validity. The upshot of this work for practitioners is that, while the case-crossover design can be useful for testing the causal null hypothesis in the presence of baseline confounders, extra caution is warranted when using the case-crossover design for point estimation of causal effects.

stat.ME

G-Net: A Deep Learning Approach to G-computation for Counterfactual Outcome Prediction Under Dynamic Treatment Regimes

Counterfactual prediction is a fundamental task in decision-making. G-computation is a method for estimating expected counterfactual outcomes under dynamic time-varying treatment strategies. Existing G-computation implementations have mostly employed classical regression models with limited capacity to capture complex temporal and nonlinear dependence structures. This paper introduces G-Net, a novel sequential deep learning framework for G-computation that can handle complex time series data while imposing minimal modeling assumptions and provide estimates of individual or population-level time varying treatment effects. We evaluate alternative G-Net implementations using realistically complex temporal simulated data obtained from CVSim, a mechanistic model of the cardiovascular system.

cs.LG

Efficient estimation of optimal regimes under a no direct effect assumption

We derive new estimators of an optimal joint testing and treatment regime under the no direct effect (NDE) assumption that a given laboratory, diagnostic, or screening test has no effect on a patient's clinical outcomes except through the effect of the test results on the choice of treatment. We model the optimal joint strategy using an optimal regime structural nested mean model (opt-SNMM). The proposed estimators are more efficient than previous estimators of the parameters of an opt-SNMM because they efficiently leverage the `no direct effect (NDE) of testing' assumption. Our methods will be of importance to decision scientists who either perform cost-benefit analyses or are tasked with the estimation of the `value of information' supplied by an expensive diagnostic test (such as an MRI to screen for lung cancer).

stat.ME