Searcharxiv⌕ Search

arXiv subjects

Karla Diaz-Ordaz

Publications and source records attributed to Karla Diaz-Ordaz.

At least 19 recordsLinked to original sources

Targeted Highly Adaptive Lasso Minimum Loss Estimation of Target Functions

We propose a Targeted Highly Adaptive Lasso for estimation of non-pathwise differentiable functional parameters such as the dose-response curve (DRC) for continuous exposure. We assume the target function lies in the $k$-th order smoothness class used to define the $k$-th order Highly Adaptive Lasso (HAL), which can be well approximated by linear spans of $k$-th order spline basis functions. We construct a projection of the true target function onto a large finite dimensional working model spanned by an initial set of $k$-th order spline basis functions, which defines a pathwise differentiable approximation of the target functional parameter. A standard TMLE is then applied with a data-adaptive initial fit, replacing the MLE targeting step with a LASSO step over HAL spline basis functions that span the target function. We prove that the resulting Targeted HAL-MLE is pointwise asymptotically normally distributed and achieves a convergence rate determined solely by the dimension and smoothness of the target function, giving dimension free rates up till $\log n$-factors. Through a simulation study for the DRC, we show that the Targeted HAL outperforms a HAL plug-in estimator in terms of bias and mean squared error. Targeted HAL offers a fully data-adaptive approach to inference on functional parameters without requiring sieve specification or parametric assumptions.

stat.ME↗

Evaluating the impact of longitudinal treatment strategies in the presence of informative monitoring and time-dependent confounding

Routinely collected data from electronic health records (EHR) provide opportunities to study effects of longitudinal treatment strategies in real-world clinical settings. A challenge presented by EHR data is that frequency of covariate monitoring differs by patient, covariate type and over time, and may be informative about a patient's health status. Many causal inference methods assume measurements of covariates are observed at a common set of regular time points. In this paper we describe and evaluate methods for estimating causal effects of longitudinal treatments on time-to-event outcomes in the presence of informative monitoring of time-dependent confounders. We show how methods based on inverse probability weighting, G-computation and longitudinal targeted maximum likelihood estimation (TMLE) can be adapted to allow for informative monitoring by incorporating monitoring indicator variables as additional time-dependent confounders. We evaluate these methods using a simulation study, comparing against more simple approaches that ignore monitoring variables. We demonstrate that ignoring monitoring can result in biased estimates of treatment effects. The methods are illustrated through an investigation into the effect of early versus delayed initiation of invasive mechanical ventilation on mortality of intensive care patients using routinely-collected data from an intensive care unit. We consider static treatment strategies such as `always treat' and `never treat' but also generalise to treatment strategies that allow for flexibility in the exact initiation time and duration of treatment.

stat.AP↗

Causal-ICM: A Data Fusion Framework For Heterogeneous Treatment Effect Estimation With Multi-Task Gaussian Processes

Bridging the gap between internal and external validity is crucial for heterogeneous treatment effect estimation. Randomised controlled trials (RCTs), favoured for their internal validity due to randomisation, often encounter challenges in generalising findings due to strict eligibility criteria. Observational studies, on the other hand, may provide stronger external validity through larger and more representative samples but can suffer from compromised internal validity due to unmeasured confounding. Motivated by these complementary characteristics, we propose a novel Bayesian nonparametric approach, Causal-ICM, leveraging multi-task Gaussian processes to integrate data from both RCTs and observational studies. In particular, we introduce a parameter that controls the degree of borrowing between the datasets and prevents the observational dataset from dominating the estimation. We propose a data-adaptive procedure for choosing the optimal value of the parameter. Causal-ICM outperforms other data fusion methods in point estimation across the covariate support of the observational study and provides principled uncertainty quantification for the estimated treatment effects. We demonstrate the robust performance of Causal-ICM in diverse scenarios through multiple simulation studies and a real-world study.

stat.ME↗

Targeted learning of heterogeneous treatment effect curves for right censored or left truncated time-to-event data

In recent years, there has been growing interest in causal machine learning estimators for quantifying subject-specific effects of a binary treatment on time-to-event outcomes. Estimation approaches have been proposed which attenuate the inherent regularisation bias in machine learning predictions, with each of these estimators addressing measured confounding, right censoring, and in some cases, left truncation. However, the existing approaches are found to exhibit suboptimal finite-sample performance, with none of the existing estimators fully leveraging the temporal structure of the data, yielding non-smooth treatment effects over time. We address these limitations by introducing surv-iTMLE, a targeted learning procedure for estimating the difference in the conditional survival probabilities under two treatments. Unlike existing estimators, surv-iTMLE accommodates both left truncation and right censoring while enforcing smoothness and boundedness of the estimated treatment effect curve over time. Through extensive simulation studies under both right censoring and left truncation scenarios, we demonstrate that surv-iTMLE outperforms existing methods in terms of bias and smoothness of time-varying effect estimates in finite samples. We then illustrate surv-iTMLE's practical utility by exploring heterogeneity in the effects of immunotherapy on survival among non-small cell lung cancer (NSCLC) patients, revealing clinically meaningful temporal patterns that existing estimators may obscure.

stat.ME↗

The super learner for time-to-event outcomes: A tutorial

Estimating risks or survival probabilities conditional on individual characteristics based on censored time-to-event data is a commonly faced task. This may be for the purpose of developing a prediction model or may be part of a wider estimation procedure, such as in causal inference. A challenge is that it is impossible to know at the outset which of a set of candidate models will provide the best risk estimates. The super learner is a powerful approach for finding the best model or combination of models ('ensemble') among a pre-specified set of candidate models or 'learners', which can include both 'statistical' models (e.g. parametric, semi-parametric models) and 'machine learning' models. Super learners for time-to-event outcomes have been developed, but the literature is technical and the full details of how these methods work and can be implemented in practice have not previously been presented in an accessible format. In this paper we provide a practical tutorial on super learner methods for time-to-event outcomes. An overview of the general steps involved in the super learner is given, followed by details of three specific implementations for time-to-event outcomes. These include the originally proposed super learner, which involves using a discrete time scale, and two more recently proposed versions of the super learner for continuous-time data. We compare the properties of the methods and provide information on how they can be implemented in R. The methods are illustrated using an open access data set and R code is provided.

stat.ME↗

Parameterising the effect of a continuous treatment using average derivative effects

The average treatment effect (ATE) is commonly used to quantify the main effect of a binary treatment on an outcome. Extensions to continuous treatments are usually based on the dose-response curve or shift interventions, but both require strong overlap conditions and the resulting curves may be difficult to summarise. We focus instead on average derivative effects (ADEs) that are scalar estimands related to infinitesimal shift interventions requiring only local overlap assumptions. ADEs, however, are rarely used in practice because their estimation usually requires estimating conditional density functions. By characterising the Riesz representers of weighted ADEs, we propose a new class of estimands that provides a unified view of weighted ADEs/ATEs when the treatment is continuous/binary. We derive the estimand in our class that minimises the nonparametric efficiency bound, thereby extending optimal weighting results from the binary treatment literature to the continuous setting. We develop efficient estimators for two weighted ADEs that avoid density estimation and are amenable to modern machine learning methods, which we evaluate in simulations and an applied analysis of Warfarin dosage effects.

math.ST↗

Variable importance measures for heterogeneous treatment effects

Motivated by applications in precision medicine and treatment effect heterogeneity, recent research has focused on estimating conditional average treatment effects (CATEs) using machine learning (ML). CATE estimates may represent complicated functions that provide little insight into the key drivers of heterogeneity. Therefore, we introduce nonparametric treatment effect variable importance measures (TE-VIMs), based on the mean-squared error (MSE) in predicting the individual treatment effect. More precisely, TE-VIMs represent the increase in MSE when variables are removed from the CATE conditioning set. We derive efficient TE-VIM estimators which can be used with any CATE estimation strategy and are amenable to ML estimation. We propose several strategies to calculate these VIMs (e.g. leave-one out, or keep-one in), using popular meta-learners for the CATE. We study the finite sample performance through a simulation study and illustrate their application using clinical trial data.

stat.ME↗

Causal machine learning for heterogeneous treatment effects in the presence of missing outcome data

When estimating heterogeneous treatment effects, missing outcome data can complicate treatment effect estimation, causing certain subgroups of the population to be poorly represented. In this work, we discuss this commonly overlooked problem and consider the impact that missing at random (MAR) outcome data has on causal machine learning estimators for the conditional average treatment effect (CATE). We propose two de-biased machine learning estimators for the CATE, the mDR-learner and mEP-learner, which address the issue of under-representation by integrating inverse probability of censoring weights into the DR-learner and EP-learner respectively. We show that under reasonable conditions, these estimators are oracle efficient, and illustrate their favorable performance through simulated data settings, comparing them to existing CATE estimators, including comparison to estimators which use common missing data techniques. We present an example of their application using the GBSG2 trial, exploring treatment effect heterogeneity when comparing hormonal therapies to non-hormonal therapies among breast cancer patients post surgery, and offer guidance on the decisions a practitioner must make when implementing these estimators.

stat.ML↗

Optimally weighted average derivative effects

Weighted average derivative effects (WADEs) are nonparametric estimands with uses in economics and causal inference. Debiased WADE estimators typically require learning the conditional mean outcome as well as a Riesz representer (RR) that characterises the requisite debiasing corrections. RR estimators for WADEs often rely on kernel estimators, introducing complicated bandwidth-dependant biases. In our work we propose a new class of RRs that are isomorphic to the class of WADEs and we derive the WADE weight that is optimal, in the sense of having minimum nonparametric efficiency bound. Our optimal WADE estimators require estimating conditional expectations only (e.g. using machine learning), thus overcoming the limitations of kernel estimators. Moreover, we connect our optimal WADE to projection parameters in partially linear models. We ascribe a causal interpretation to WADE and projection parameters in terms of so-called incremental effects. We propose efficient estimators for two WADE estimands in our class, which we evaluate in a numerical experiment and use to determine the effect of Warfarin dose on blood clotting function.

stat.ME↗

Explainable AI for survival analysis: a median-SHAP approach

With the adoption of machine learning into routine clinical practice comes the need for Explainable AI methods tailored to medical applications. Shapley values have sparked wide interest for locally explaining models. Here, we demonstrate their interpretation strongly depends on both the summary statistic and the estimator for it, which in turn define what we identify as an 'anchor point'. We show that the convention of using a mean anchor point may generate misleading interpretations for survival analysis and introduce median-SHAP, a method for explaining black-box models predicting individual survival times.

cs.LG↗

PWSHAP: A Path-Wise Explanation Model for Targeted Variables

Predictive black-box models can exhibit high accuracy but their opaque nature hinders their uptake in safety-critical deployment environments. Explanation methods (XAI) can provide confidence for decision-making through increased transparency. However, existing XAI methods are not tailored towards models in sensitive domains where one predictor is of special interest, such as a treatment effect in a clinical model, or ethnicity in policy models. We introduce Path-Wise Shapley effects (PWSHAP), a framework for assessing the targeted effect of a binary (e.g.~treatment) variable from a complex outcome model. Our approach augments the predictive model with a user-defined directed acyclic graph (DAG). The method then uses the graph alongside on-manifold Shapley values to identify effects along causal pathways whilst maintaining robustness to adversarial attacks. We establish error bounds for the identified path-wise Shapley effects and for Shapley values. We show PWSHAP can perform local bias and mediation analyses with faithfulness to the model. Further, if the targeted variable is randomised we can quantify local effect modification. We demonstrate the resolution, interpretability, and true locality of our approach on examples and a real-world experiment.

stat.ML↗

Challenges and Opportunities of Shapley values in a Clinical Context

With the adoption of machine learning-based solutions in routine clinical practice, the need for reliable interpretability tools has become pressing. Shapley values provide local explanations. The method gained popularity in recent years. Here, we reveal current misconceptions about the ``true to the data'' or ``true to the model'' trade-off and demonstrate its importance in a clinical context. We show that the interpretation of Shapley values, which strongly depends on the choice of a reference distribution for modeling feature removal, is often misunderstood. We further advocate that for applications in medicine, the reference distribution should be tailored to the underlying clinical question. Finally, we advise on the right reference distributions for specific medical use cases.

stat.ME↗

Dynamic Updating of Clinical Survival Prediction Models in a Rapidly Changing Environment

Over time, the performance of clinical prediction models may deteriorate due to changes in clinical management, data quality, disease risk and/or patient mix. Such prediction models must be updated in order to remain useful. Here, we investigate methods for discrete and dynamic model updating of clinical survival prediction models based on refitting, recalibration and Bayesian updating. In contrast to discrete or one-time updating, dynamic updating refers to a process in which a prediction model is repeatedly updated with new data. Motivated by infectious disease settings, our focus was on model performance in rapidly changing environments. We first compared the methods using a simulation study. We simulated scenarios with changing survival rates, the introduction of a new treatment and predictors of survival that are rare in the population. Next, the updating strategies were applied to patient data from the QResearch database, an electronic health records database from general practices in the UK, to study the updating of a model for predicting 70-day covid-19 related mortality. We found that a dynamic updating process outperformed one-time discrete updating in the simulations. Bayesian dynamic updating has the advantages of making use of knowledge from previous updates and requiring less data compared to refitting.

stat.ME↗

Demystifying statistical learning based on efficient influence functions

Evaluation of treatment effects and more general estimands is typically achieved via parametric modelling, which is unsatisfactory since model misspecification is likely. Data-adaptive model building (e.g. statistical/machine learning) is commonly employed to reduce the risk of misspecification. Naive use of such methods, however, delivers estimators whose bias may shrink too slowly with sample size for inferential methods to perform well, including those based on the bootstrap. Bias arises because standard data-adaptive methods are tuned towards minimal prediction error as opposed to e.g. minimal MSE in the estimator. This may cause excess variability that is difficult to acknowledge, due to the complexity of such strategies. Building on results from non-parametric statistics, targeted learning and debiased machine learning overcome these problems by constructing estimators using the estimand's efficient influence function under the non-parametric model. These increasingly popular methodologies typically assume that the efficient influence function is given, or that the reader is familiar with its derivation. In this paper, we focus on derivation of the efficient influence function and explain how it may be used to construct statistical/machine-learning-based estimators. We discuss the requisite conditions for these estimators to perform well and use diverse examples to convey the broad applicability of the theory.

math.ST↗

On Locality of Local Explanation Models

Shapley values provide model agnostic feature attributions for model outcome at a particular instance by simulating feature absence under a global population distribution. The use of a global population can lead to potentially misleading results when local model behaviour is of interest. Hence we consider the formulation of neighbourhood reference distributions that improve the local interpretability of Shapley values. By doing so, we find that the Nadaraya-Watson estimator, a well-studied kernel regressor, can be expressed as a self-normalised importance sampling estimator. Empirically, we observe that Neighbourhood Shapley values identify meaningful sparse feature relevance attributions that provide insight into local model behaviour, complimenting conventional Shapley analysis. They also increase on-manifold explainability and robustness to the construction of adversarial classifiers.

cs.LG↗

Robust Inference for Mediated Effects in Partially Linear Models

We consider mediated effects of an exposure, X on an outcome, Y, via a mediator, M, under no unmeasured confounding assumptions in the setting where models for the conditional expectation of the mediator and outcome are partially linear. We propose G-estimators for the direct and indirect effect and demonstrate consistent asymptotic normality for indirect effects when models for the conditional means of M, or X and Y are correctly specified, and for direct effects, when models for the conditional means of Y, or X and M are correct. This marks an improvement, in this particular setting, over previous `triple' robust methods, which do not assume partially linear mean models. Testing of the no-mediation hypothesis is inherently problematic due to the composite nature of the test (either X has no effect on M or M no effect on Y), leading to low power when both effect sizes are small. We use Generalized Methods of Moments (GMM) results to construct a new score testing framework, which includes as special cases the no-mediation and the no-direct-effect hypotheses. The proposed tests rely on an orthogonal estimation strategy for estimating nuisance parameters. Simulations show that the GMM based tests perform better in terms of power and small sample performance compared with traditional tests in the partially linear setting, with drastic improvement under model misspecification. New methods are illustrated in a mediation analysis of data from the COPERS trial, a randomized trial investigating the effect of a non-pharmacological intervention of patients suffering from chronic pain. An accompanying R package implementing these methods can be found at github.com/ohines/plmed.

stat.ME↗

How to estimate the association between change in a risk factor and a health outcome?

Estimating the effect of a change in a particular risk factor and a chronic disease requires information on the risk factor from two time points; the enrolment and the first follow-up. When using observational data to study the effect of such an exposure (change in risk factor) extra complications arise, namely (i) when is time zero? and (ii) which information on confounders should we account for in this type of analysis? From enrolment or the 1st follow-up? Or from both?. The combination of these questions has proven to be very challenging. Researchers have applied different methodologies with mixed success, because the different choices made when answering these questions induce systematic bias. Here we review these methodologies and highlight the sources of bias in each type of analysis. We discuss the advantages and the limitations of each method ending by making our recommendations on the analysis plan.

stat.ME↗

Missing binary outcomes under covariate dependent missingness in cluster randomised trials

Missing outcomes are a commonly occurring problem for cluster randomised trials, which can lead to biased and inefficient inference if ignored or handled inappropriately. Two approaches for analysing such trials are cluster-level analysis and individual-level analysis. In this study, we assessed the performance of unadjusted cluster-level analysis, baseline covariate adjusted cluster-level analysis, random effects logistic regression (RELR) and generalised estimating equations (GEE) when binary outcomes are missing under a baseline covariate dependent missingness mechanism. Missing outcomes were handled using complete records analysis (CRA) and multilevel multiple imputation (MMI). We analytically show that cluster-level analyses for estimating risk ratio (RR) using complete records are valid if the true data generating model has log link and the intervention groups have the same missingness mechanism and the same covariate effect in the outcome model. We performed a simulation study considering four different scenarios, depending on whether the missingness mechanisms are the same or different between the intervention groups and whether there is an interaction between intervention group and baseline covariate in the outcome model. Based on the simulation study and analytical results, we give guidance on the conditions under which each approach is valid.

stat.ME↗