SearcharxivSearch

arXiv subjects

Jared D. Huling

Publications and source records attributed to Jared D. Huling.

At least 19 recordsLinked to original sources

Nonparametric Efficient Estimation of Dynamic Treatment Regimes with Competing Risks

Many observational studies evaluate the risks and benefits of time-varying interventions. In these longitudinal settings, the primary outcome of interest is often a clinical event that is precluded by mortality, which acts as a competing risk. Evaluating dynamic treatment regimes (DTRs) is complicated by time-varying confounding, right-censoring, and progressive sample size reduction over time, for which traditional methods such as the g-formula or inverse probability weighting may be misspecified or highly unstable. In this work, we develop the nonparametric efficiency theory and derive the efficient influence function for the cumulative incidence of an event under a DTR in the presence of a competing risk and right-censoring. Building on this result, we propose a sequentially doubly robust estimator that accommodates flexible machine learning for nuisance estimation while retaining root-n consistency and asymptotic normality. We illustrate the utility of our framework by estimating long-term cumulative incidence of clinical fractures under different bisphosphonate "drug holiday" regimes for osteoporosis.

stat.ME

Causal Inference for Functional Treatments with Stochastic Policies

Wearable devices can accurately measure human behavior, providing a unique opportunity to understand how behavior impacts health. Recent studies leveraging functional regression methods have found a strong relationship between accelerometer-collected physical activity and mortality. However, to determine if physical activity patterns impact mortality it is necessary to understand the causal effects of policies for physical activity, i.e., a function-valued treatment. Functional treatments present several challenges for causal effect estimation: 1) defining a scientifically meaningful estimand that reflects real-world policies and satisfies positivity is nontrivial; and 2) the potential for temporal confounding over continuous time. To address these, we propose stochastic policies for functional treatments that allow estimation of causal effects of changing the treatment distribution without requiring a positivity assumption. We develop a novel method for such that modifies the treatment through a single basis function chosen by the analyst, allowing for clear control over treatment modification and temporal confounding feedback. We show asymptotic normality of our estimators and that they exhibit rate double robustness. We apply our methods to the National Health and Nutrition Examination Survey to determine the causal effect of increasing physical activity over three-hour periods on mortality.

stat.ME

Omitted-Variable Sensitivity Analysis for Generalizing Randomized Trials

Randomized controlled trials (RCTs) yield internally valid causal effect estimates, but generalizing these results to target populations with different characteristics requires an untestable selection ignorability assumption: conditional on observed covariates, trial participation must be independent of potential outcomes. This assumption fails when unobserved effect modifiers are distributed differently between trial and target populations. We develop a sensitivity analysis framework for trial generalization grounded in omitted variable bias (OVB). Our key theoretical contribution is an exact decomposition showing that external-validity bias equals moderation strength $\times$ moderator imbalance: (i) how strongly an unobserved variable shifts the treatment effect, times (ii) how differently that variable is distributed across populations after covariate adjustment. We introduce scale-free sensitivity parameters based on partial $R^2$ values, enabling closed-form bounds and benchmarking against observed covariates -- practitioners can assess whether conclusions would change if an unobserved moderator were "as strong as" a particular observed variable. Simulations demonstrate that our bounds achieve nominal coverage and remain conservative under model misspecification, while comparisons with alternative sensitivity frameworks highlight the interpretive advantages of the OVB decomposition.

stat.ME

Heterogeneous readmission prediction with hierarchical effect decomposition and regularization

Accurately predicting hospital readmission risks using electronic health records (EHRs) is critical for effective patient management and healthcare resource allocation. Patient populations in health systems are highly heterogeneous across different primary diagnoses, necessitating tailored yet interpretable prediction models. We propose a hierarchical modeling framework incorporating hierarchical nested re-parameterization and structured regularization methods, which we call hierNest. Specifically, our approach leverages the inherent hierarchical structure present in primary diagnoses and groupings of these diagnoses into major diagnostic categories. Our methodology facilitates information borrowing across related patient subgroups and preserves interpretability at different hierarchical levels. Simulation studies demonstrate superior predictive accuracy of the proposed method, particularly with small subgroup sample sizes and varying degrees of hierarchical effects. We apply our methods to a large EHR dataset comprising Medicare patients.

stat.ME

Improving RCT-Based CATE Estimation Under Covariate Mismatch via Double Calibration

We develop estimators that improve precision of heterogeneous treatment effect estimates that allow borrowing information from observational studies when the available covariates in each data source do not perfectly match. Standard data-borrowing methods often assume perfectly matched covariates. We propose MR-OSCAR, an RCT-calibrated, two-stage estimation approach that first predicts the trial-missing variables using the observational data via imputation and then calibrates observational outcome predictions to the randomized trial, preserving the causal contrast, unlike the results for generalization, where imputation does not improve performance. Our theory gives finite-sample guarantees with a transparent error decomposition including an imputation error that shrinks as the observational mapping becomes more predictable. Simulations show that imputation almost always outperforms naively using only the shared covariates and clarifies when borrowing helps (strong predictability of the missing block, moderate trial size) and when it does not (poor predictability or dominant trial-only moderators). We motivate the approach with the Greenlight Plus trial on early childhood obesity and outline a forthcoming EHR analysis at Vanderbilt, highlighting the use of our method in common scenarios where data do not perfectly align.

stat.ME

Estimation of heterogeneous principal effects under principal ignorability

We study estimation and inference for heterogeneous principal causal effects with binary treatments and binary intermediate variables. Principal causal effects are subgroup effects within strata defined by potential values of an intermediate variable, including effects among compliers. We propose a framework for estimating and forming pointwise confidence intervals for heterogeneous principal causal effects under the principal ignorability assumption. Several estimators are developed, and their robustness properties are characterized: one estimator is doubly robust, whereas the other two attain intermediate robustness between double and triple robustness; in contrast, principal causal effects can be estimated in a triply robust manner only. We establish large-sample theory under nonparametric smoothness conditions and analyze the bias contributions of each approach, providing insight into performance beyond the smooth setting, including in high-dimensional regimes. Camden Coalition hotspotting randomized trial are used to illustrate the methods by estimating heterogeneous complier effects.

stat.ME

Sharp Bounds for Treatment Effect Generalization under Outcome Distribution Shift

Generalizing treatment effects from a randomized trial to a target population requires the assumption that potential outcome distributions are invariant across populations after conditioning on observed covariates. This assumption fails when unmeasured effect modifiers are distributed differently between trial participants and the target population. We develop a sensitivity analysis framework that bounds how much conclusions can change when this transportability assumption is violated. Our approach constrains the likelihood ratio between target and trial outcome densities by a scalar parameter $\Lambda \geq 1$, with $\Lambda = 1$ recovering standard transportability. For each $\Lambda$, we derive sharp bounds on the target average treatment effect -- the tightest interval guaranteed to contain the true effect under all data-generating processes compatible with the observed data and the sensitivity model. We show that the optimal likelihood ratios have a simple threshold structure, leading to a closed-form greedy algorithm that requires only sorting trial outcomes and redistributing probability mass. The resulting estimator runs in $O(n \log n)$ time and is consistent under standard regularity conditions. Simulations demonstrate that our bounds achieve nominal coverage when the true outcome shift falls within the specified $\Lambda$, provide substantially tighter intervals than worst-case bounds, and remain informative across a range of realistic violations of transportability.

stat.ME

Estimating causal effects of functional treatments with modified functional treatment policies

Functional data are increasingly prevalent in biomedical research. While functional data analysis has been established for decades, causal inference with functional treatments remains largely unexplored. Existing methods typically focus on estimating the causal average dose response functional (ADRF), which requires strong positivity assumptions and offers limited interpretability. In this work, we target a new causal estimand, the modified functional treatment policy (MFTP), which focuses on estimating the average potential outcome when each individual slightly modifies their treatment trajectory from the observed one. A major challenge for this new estimand is the need to define an average over an infinite-dimensional object with no density. By proposing a novel definition of the population average over a functional variable using a functional principal component analysis (FPCA) decomposition, we establish the causal identifiability of the MFTP estimand. We further derive outcome regression, inverse probability weighting, and doubly robust estimators for the MFTP, and provide theoretical guarantees under mild regularity conditions. The proposed estimators are validated through extensive simulation studies. Applying our MFTP framework to the National Health and Nutrition Examination Survey (NHANES) accelerometer data, we estimate the causal effects of reducing disruptive nighttime activity and low-activity duration on all-cause mortality.

stat.ME

Data adaptive covariate balancing for causal effect estimation for high dimensional data

A key challenge in estimating causal effects from observational data is handling confounding and is commonly achieved through weighting methods that balance distribution of covariates between treatment and control groups. Weighting approaches can be classified by whether weights are estimated using parametric or nonparametric methods, and by whether the model relies on modeling and inverting the propensity score or directly estimates weights to achieve distributional balance by minimizing a measure of dissimilarity between groups. Parametric methods, both for propensity score modeling and direct balancing, are prone to model misspecification. In addition, balancing approaches often suffer from the curse of dimensionality, as they assign equal importance to all covariates, thus potentially de-emphasizing true confounders. Several methods, such as the outcome adaptive lasso, attempt to mitigate this issue through variable selection, but are parametric and focus on propensity score estimation rather than direct balancing. In this paper, we propose a nonparametric direct balancing approach that uses random forests to adaptively emphasize confounders. Our method jointly models treatment and outcome using random forests, allowing the data to identify covariates that influence both processes. We construct a similarity measure, defined by the proportion of trees in which two observations fall into the same leaf node, yielding a distance between treatment and control distributions that is sensitive to relevant covariates and captures the structure of confounding. Under suitable assumptions, we show that the resulting weights converge to normalized inverse propensity scores in the L2 norm and provide consistent treatment effect estimates. We demonstrate the effectiveness of our approach through extensive simulations and an application to a real dataset.

stat.ME

Partially Retargeted Balancing Weights for Causal Effect Estimation Under Positivity Violations

Positivity violations, which occur when some subgroups either always or never receive a treatment of interest, pose significant challenges for causal effect estimation with observational data. Recent balancing weight methods have proved to be highly effective in confounding control, however their utility is diminished in the presence of positivity violations, resulting in bias and excess variance. Approaches that deal with positivity violations, on the other hand, work by targeting a modified estimand that may be misaligned with the original research question. To address these challenges, we propose a novel balancing weights approach, which mitigates positivity violations while attempting to retain the original estimand by a targeted relaxation of the balancing constraints. Our proposed weighted estimator is consistent for the original estimand when either 1) the implied propensity score model is correct; or 2) all treatment effect modifiers are balanced to the target population. When these conditions do not hold, our estimator is consistent for a slightly modified treatment effect estimand. Furthermore, our proposed weighted estimator has reduced asymptotic variance when positivity does not hold. We evaluate our approach through applications to synthetic data, an observational study, and when transporting a treatment effect from a randomized trial.

stat.ME

Exploring the effects of mechanical ventilator settings with modified vector-valued treatment policies

Mechanical ventilation is critical for managing respiratory failure, but inappropriate ventilator settings can lead to ventilator-induced lung injury (VILI), increasing patient morbidity and mortality. Evaluating the causal impact of ventilator settings is challenging due to the complex interplay of multiple treatment variables and strong confounding due to ventilator guidelines. In this paper, we propose a modified vector-valued treatment policy (MVTP) framework coupled with energy balancing weights to estimate causal effects involving multiple continuous ventilator parameters simultaneously in addition to sensitivity analysis to unmeasured confounding. Our approach mitigates common challenges in causal inference for vector-valued treatments, such as infeasible treatment combinations, stringent positivity assumptions, and interpretability concerns. Using the MIMIC-III database, our analyses suggest that equal reductions in the total power of ventilation (i.e., the mechanical power) through different ventilator parameters result in different expected patient outcomes. Specifically, lowering airway pressures may yield greater reductions in patient mortality compared to proportional adjustments of tidal volume alone. Moreover, controlling for respiratory-system compliance and minute ventilation, we found a significant benefit of reducing driving pressure in patients with acute respiratory distress syndrome (ARDS). Our analyses help shed light on the contributors to VILI.

stat.ME

Heterogeneity-Aware Regression with Nonparametric Estimation and Structured Selection for Hospital Readmission Prediction

Readmission prediction is a critical but challenging clinical task, as the inherent relationship between high-dimensional covariates and readmission is complex and heterogeneous. Despite this complexity, models should be interpretable to aid clinicians in understanding an individual's risk prediction. Readmissions are often heterogeneous, as individuals hospitalized for different reasons, particularly across distinct clinical diagnosis groups, exhibit materially different subsequent risks of readmission. To enable flexible yet interpretable modeling that accounts for patient heterogeneity, we propose a novel hierarchical-group structure kernel that uses sparsity-inducing kernel summation for variable selection. Specifically, we design group-specific kernels that vary across clinical groups, with the degree of variation governed by the underlying heterogeneity in readmission risk; when heterogeneity is minimal, the group-specific kernels naturally align, approaching a shared structure across groups. Additionally, by allowing variable importance to adapt across interactions, our approach enables more precise characterization of higher-order effects, improving upon existing methods that capture nonlinear and higher-order interactions via functional ANOVA. Extensive simulations and a hematologic readmission dataset (n=18,096) demonstrate superior performance across subgroups of patients (AUROC, PRAUC) over the lasso and XGBoost. Additionally, our model provides interpretable insights into variable importance and group heterogeneity.

stat.ME

A Unified Framework for Causal Estimand Selection

Estimating the causal effect of a treatment or health policy with observational data can be challenging due to an imbalance of and a lack of overlap between treated and control covariate distributions. In the presence of limited overlap, researchers choose between 1) methods (e.g., inverse probability weighting) that imply traditional estimands but whose estimators are at risk of considerable bias and variance; and 2) methods (e.g., overlap weighting) which imply a different estimand, thereby modifying the target population to reduce variance. We propose a framework for navigating the tradeoffs between variance and bias due to imbalance and lack of overlap and the targeting of the estimand of scientific interest. We introduce a bias decomposition that encapsulates bias due to 1) the statistical bias of the estimator; and 2) estimand mismatch, i.e., deviation from the population of interest. We propose two design-based metrics and an estimand selection procedure that help illustrate the tradeoffs between these sources of bias and variance of the resulting estimators. Our procedure allows analysts to incorporate their domain-specific preference for preservation of the original research population versus reduction of statistical bias. We demonstrate how to select an estimand based on these preferences with an application to right heart catheterization data.

stat.ME

Transportability of Principal Causal Effects

Recent research in causal inference has made important progress in addressing challenges to the external validity of trial findings. Such methods weight trial participant data to more closely resemble the distribution of effect-modifying covariates in a well-defined target population. In the presence of participant non-adherence to study medication, these methods effectively transport an intention-to-treat effect that averages over heterogeneous compliance behaviors. In this paper, we develop a principal stratification framework to identify causal effects conditioning on both compliance behavior and membership in the target population. We also develop non-parametric efficiency theory for and construct efficient estimators of such "transported" principal causal effects and characterize their finite-sample performance in simulation experiments. While this work focuses on treatment non-adherence, the framework is applicable to a broad class of estimands that target effects in clinically-relevant, possibly latent subsets of a target population.

stat.ME

A reluctant additive model framework for interpretable nonlinear individualized treatment rules

Individualized treatment rules (ITRs) for treatment recommendation is an important topic for precision medicine as not all beneficial treatments work well for all individuals. Interpretability is a desirable property of ITRs, as it helps practitioners make sense of treatment decisions, yet there is a need for ITRs to be flexible to effectively model complex biomedical data for treatment decision making. Many ITR approaches either focus on linear ITRs, which may perform poorly when true optimal ITRs are nonlinear, or black-box nonlinear ITRs, which may be hard to interpret and can be overly complex. This dilemma indicates a tension between interpretability and accuracy of treatment decisions. Here we propose an additive model-based nonlinear ITR learning method that balances interpretability and flexibility of the ITR. Our approach aims to strike this balance by allowing both linear and nonlinear terms of the covariates in the final ITR. Our approach is parsimonious in that the nonlinear term is included in the final ITR only when it substantially improves the ITR performance. To prevent overfitting, we combine cross-fitting and a specialized information criterion for model selection. Through extensive simulations, we show that our methods are data-adaptive to the degree of nonlinearity and can favorably balance ITR interpretability and flexibility. We further demonstrate the robust performance of our methods with an application to a cancer drug sensitive study.

stat.ME

Modified treatment policy effect estimation with weighted energy distance

The causal effects of continuous treatments are often characterized through the average dose response function, which is challenging to estimate from observational data due to confounding and positivity violations. Modified treatment policies (MTPs) are an alternative approach that aim to assess the effect of a modification to observed treatment values and work under relaxed assumptions. Estimators for MTPs generally focus on estimating the conditional density of treatment given covariates and using it to construct weights. However, weighting using conditional density models has well-documented challenges. Further, MTPs with larger treatment modifications have stronger confounding and no tools exist to help choose an appropriate modification magnitude. This paper investigates the role of weights for MTPs showing that to control confounding, weights should balance the weighted data to an unobserved hypothetical target population that can be characterized with observed data. Leveraging this insight, we present a versatile set of tools to enhance estimation for MTPs. We introduce a distance that measures imbalance of covariate distributions under the MTP and use it to develop new weighting methods and tools to aid in the estimation of MTPs. Using our methods we study the effect of mechanical power of ventilation on in-hospital mortality.

stat.ME

Improving Precision of RCT-Based CATE Estimation using Data Borrowing with Double Calibration

Understanding how treatment effects vary across patient characteristics is essential for personalized medicine, yet randomized controlled trials (RCTs) are often underpowered to detect heterogeneous treatment effects (HTEs). We propose a framework that improves the efficiency of conditional average treatment effect (CATE) estimation in RCTs by leveraging large observational studies (OS) while preserving RCT unbiasedness. Framing CATE estimation as a supervised learning problem, we show that estimation variance is minimized using the counterfactual mean outcome (CMO) as an augmentation function. We derive finite-sample error bounds and give conditions under which OS data improves CMO estimation, and thus CATE efficiency, even under confounding in the OS or outcome distribution shift between populations. We introduce R-OSCAR (Robust Observational Studies for CMO-Augmented RCT), a two-stage estimator that calibrates OS outcome predictions to the RCT population and corrects residual bias through regularized regression. For any OS-derived nuisance, R-OSCAR is consistent for the RCT-population CATE, and is efficient relative to RCT-only estimators when the RCT-OS outcome mean discrepancy is estimable from the RCT at lower complexity than the full RCT outcome model. A cross-fitted RCT diagnostic determines, from observable data alone, whether borrowing from a given OS is supported. Simulations show R-OSCAR can reduce the RCT sample size needed for HTE detection by up to 75%, while remaining robust to misspecification. We validate on two case studies: a semi-synthetic analysis of the Tennessee STAR study with constructed observational confounding, and the Greenlight Plus pediatric-obesity trial linked with external electronic-health-record controls, where borrowing improves control-arm estimation for small trials and the diagnostic certifies it only where the records cover the trial population.

stat.ME

Doubly structured sparsity for grouped multivariate responses with application to functional outcome score modeling

This work is motivated by the need to accurately model a vector of responses related to pediatric functional status using administrative health data from inpatient rehabilitation visits. The components of the responses have known and structured interrelationships. To make use of these relationships in modeling, we develop a two-pronged regularization approach to borrow information across the responses. The first component of our approach encourages joint selection of the effects of each variable across possibly overlapping groups related responses and the second component encourages shrinkage of effects towards each other for related responses. As the responses in our motivating study are not normally-distributed, our approach does not rely on an assumption of multivariate normality of the responses. We show that with an adaptive version of our penalty, our approach results in the same asymptotic distribution of estimates as if we had known in advance which variables were non-zero and which variables have the same effects across some outcomes. We demonstrate the performance of our method in extensive numerical studies and in an application in the prediction of functional status of pediatric patients using administrative health data in a population of children with neurological injury or illness at a large children's hospital.

stat.ME