SearcharxivSearch

arXiv subjects

Carol Gao

Publications and source records attributed to Carol Gao.

4 recordsLinked to original sources

Instrumental Variables with Time-Varying Exposure: Dynamic Effects of Revascularization on Quality of Life

This paper develops instrumental variables (IV) estimators for dynamic causal effects in randomized trials with imperfect compliance. These methods are applied to a randomized trial that assigned patients with ischemic heart disease to either an invasive treatment arm centered on revascularization or a control group meant to receive non-invasive medical therapy. As is common in such ``strategy trials,'' many participants assigned to treatment remained untreated while many assigned to control crossed over into treatment. Protocol non-compliance causes ITT estimates to diverge from the effect of treatment received, while conventional per-protocol analyses that condition on treatment received are compromised by selection bias. Extending the static potential-outcomes IV framework, the methods here identify average causal effects of treatment for dynamic compliers, the set of trial participants who comply with trial protocol at different follow-up horizons. IV estimates of revascularization effects on compliers' quality of life are markedly larger and more sustained than previously reported ITT and per-protocol estimates. We also show how to estimate average characteristics and marginal potential outcome means for dynamic compliers. These results are used to explain confounding in as-treated per-protocol estimates.

econ.EM

Robustifying and Selecting Cohort-Appropriate Prognostic Models under Distributional Shifts

External validation is widely regarded as the gold standard for prognostic model evaluation. In this study, we challenge the assumption that successful external calibration guarantees model generalizability and propose two complementary strategies to improve transportability of prognostic models across cohorts. Using six real-world surgical cohorts from tertiary academic centers, we tested whether successful external calibration depends largely on similarity in covariates and outcomes between training and validation cohorts, quantified using Kullback-Leibler (KL) divergence, with calibration assessed by the Integrated Calibration Index (ICI). From the model-developer's perspective, we trained the "best-on-average" prognostic model by tuning toward a meta-analysis-derived covariate and outcome distribution as an approximation of the broader target population. From the end-user perspective, we proposed a simple measure for cohort outcome similarity to identify, among published models, the one most suitable for a given target cohort in terms of both calibration and clinical utility. External calibration worsened as distributional mismatch increased. Higher KL divergence was associated with higher ICI in both surgery-alone (Spearman $ρ=0.614$, $p=0.004$) and surgery + adjuvant chemotherapy cohorts (Spearman $ρ=0.738$, $p<0.001$). Meta-analysis-informed weighting improved calibration in most settings without materially affecting discrimination, with the clearest benefit when evaluated on the aggregated external population ($p=0.037$). Models developed in more similar cohorts achieved lower ICI in surgery-alone (Spearman $ρ=0.803$, $p<0.001$) and surgery + adjuvant chemotherapy cohorts (Spearman $ρ=0.737$, $p<0.001$), and provided greater clinical utility on DCA.

stat.ME

Should We Relax Stability in Matching Markets?

Centralized assignment markets have historically relied on Deferred-Acceptance (DA) algorithms, which do not incorporate multiple objectives into the assignment. In this work, we propose an optimization-based many-to-one assignment algorithm that explores the trade-offs between minimizing the number of blocking pairs in the match and other important objectives. In order to scale to high-dimensional problems, we develop an algorithm using inverse optimization to obtain the optimal cost vector that implicitly optimizes for stability. This is empirically tested on two application areas for which the DA algorithm is widely used: school assignment and medical residency match. Computational tests on a simulated Boston Public Schools (BPS) match show that this method effectively reduces transportation cost and increases number of students receiving an offer in the first round of match at the expense of a small percentage of blocking pairs. Similar improvement in the number of couples matched to the same location is observed in a synthetic residency match.

math.OC

The R.O.A.D. to clinical trial emulation

Observational studies provide the only evidence on the effectiveness of interventions when randomized controlled trials (RCTs) are impractical due to cost, ethical concerns, or time constraints. While many methodologies aim to draw causal inferences from observational data, there is a growing trend to model observational study designs after RCTs, a strategy known as "target trial emulation." Despite its potential, causal inference through target trial emulation cannot fully address the confounding bias in real-world data due to the lack of randomization. In this work, we present a novel framework for target trial emulation that aims to overcome several key limitations, including confounding bias. The framework proceeds as follows: First, we apply the eligibility criteria of a specific trial to an observational cohort. We then "correct" this cohort by extracting a subset that matches both the distribution of covariates and the baseline prognosis of the control group in the target RCT. Next, we address unmeasured confounding by adjusting the prognosis estimates of the treated group to align with those observed in the trial. Following trial emulation, we go a step further by leveraging the emulated cohort to train optimal decision trees, to identify subgroups of patients with heterogeneity in treatment effects (HTE). The absence of confounding is verified using two external models, and the validity of the treatment recommendations is independently confirmed by the team responsible for the original trial we emulate. To our knowledge, this is the first framework to successfully address both observed and unobserved confounding, a challenge that has historically limited the use of randomized trial emulation and causal inference. Additionally, our framework holds promise in advancing precision medicine by identifying patient subgroups that benefit most from specific treatments.

stat.AP