Searcharxiv⌕ Search

arXiv subjects

Rachael K. Ross

Publications and source records attributed to Rachael K. Ross.

4 recordsLinked to original sources

Constructing targeted minimum loss/maximum likelihood estimators: a simple illustration to build intuition

Use of machine learning to estimate nuisance functions (e.g. outcomes models, propensity score models) in estimators used in causal inference is increasingly common, as it can mitigate bias due to model misspecification. However, it can be challenging to achieve valid inference (e.g., estimate valid confidence intervals). The efficient influence function (EIF) provides a recipe to go from a statistical estimand relevant to our causal question, to an estimator that can validly incorporate machine learning. Our companion paper, Renson et al. 2025 (arXiv:2502.05363), provides a thorough but approachable description of the EIF, along with a guide through the steps to go from a unique statistical estimand to development of one type of EIF-based estimator, the so-called one-step estimator. Another commonly used estimator based on the EIF is the targeted maximum likelihood/minimum loss estimator (TMLE). Construction of TMLEs is well-discussed in the statistical literature, but there remains a gap in translation to a more applied audience. In this letter, which supplements Renson et al., we provide a more accessible illustration of how to construct a TMLE.

stat.ME↗

Transporting results from a trial to an external target population when trial participation impacts adherence

Randomized clinical trials are considered the gold standard for informing treatment guidelines, but results may not generalize to real-world populations. Generalizability is hindered by distributional differences in baseline covariates and treatment-outcome mediators. Approaches to address differences in covariates are well established, but approaches to address differences in mediators are more limited. Here we consider the setting where trial activities that differ from usual care settings (e.g., monetary compensation, follow-up visits frequency) affect treatment adherence. When treatment and adherence data are unavailable for the real-world target population, we cannot identify the mean outcome under a specific treatment assignment (i.e., mean potential outcome) in the target. Therefore, we propose a sensitivity analysis in which a parameter for the relative difference in adherence to a specific treatment between the trial and the target, possibly conditional on covariates, must be specified. We discuss options for specification of the sensitivity analysis parameter based on external knowledge including setting a range to estimate bounds or specifying a probability distribution from which to repeatedly draw parameter values (i.e., use Monte Carlo sampling). We introduce two estimators for the mean counterfactual outcome in the target that incorporates this sensitivity parameter, a plug-in estimator and a one-step estimator that is double robust and supports the use of machine learning for estimating nuisance models. Finally, we apply the proposed approach to the motivating application where we transport the risk of relapse under two different medications for the treatment of opioid use disorder from a trial to a real-world population.

stat.ME↗

Pulling back the curtain: the road from statistical estimand to machine-learning based estimator for epidemiologists (no wizard required)

Epidemiologists increasingly use causal inference methods that rely on machine learning, as these approaches can relax unnecessary model specification assumptions. While deriving and studying asymptotic properties of such estimators is a task usually associated with statisticians, it is useful for epidemiologists to understand the steps involved, as epidemiologists are often at the forefront of defining important new research questions and translating them into new parameters to be estimated. In this paper, our goal was to provide a relatively accessible guide through the process of (i) deriving an estimator based on the so-called efficient influence function (which we define and explain), and (ii) showing such an estimator's ability to validly incorporate machine learning, by demonstrating the so-called rate double robustness property. The derivations in this paper rely mainly on algebra and some foundational results from statistical inference, which are explained.

stat.ME↗

Double Robust Variance Estimation with Parametric Working Models

Doubly robust estimators have gained popularity in the field of causal inference due to their ability to provide consistent point estimates when either an outcome or exposure model is correctly specified. However, for nonrandomized exposures the influence function based variance estimator frequently used with doubly robust estimators of the average causal effect is only consistent when both working models (i.e., outcome and exposure models) are correctly specified. Here, the empirical sandwich variance estimator and the nonparametric bootstrap are demonstrated to be doubly robust variance estimators. That is, they are expected to provide valid estimates of the variance leading to nominal confidence interval coverage when only one working model is correctly specified. Simulation studies illustrate the properties of the influence function based, empirical sandwich, and nonparametric bootstrap variance estimators in the setting where parametric working models are assumed. Estimators are applied to data from the Improving Pregnancy Outcomes with Progesterone (IPOP) study to estimate the effect of maternal anemia on birth weight among women with HIV.

stat.ME↗