SearcharxivSearch

arXiv subjects

Erin E Gabriel

Publications and source records attributed to Erin E Gabriel.

12 recordsLinked to original sources

Propensity weighting plus adjustment in proportional hazards model is not doubly robust

Recently, it has become common for applied works to combine commonly used survival analysis modeling methods, such as the multivariable Cox model and propensity score weighting, with the intention of forming a doubly robust estimator of an exposure effect hazard ratio that is unbiased in large samples when either the Cox model or the propensity score model is correctly specified. This combination does not, in general, produce a doubly robust estimator, even after regression standardization, when there is truly a causal effect. We demonstrate via simulation this lack of double robustness for the semiparametric Cox model, the Weibull proportional hazards model, and a simple proportional hazards flexible parametric model, with both the latter models fit via maximum likelihood. We provide a novel proof that the combination of propensity score weighting and a proportional hazards survival model, fit either via full or partial likelihood, is consistent under the null of no causal effect of the exposure on the outcome under particular censoring mechanisms if either the propensity score or the outcome model is correctly specified and contains all confounders. Given our results suggesting that double robustness only exists under the null, we outline two simple alternative estimators that are doubly robust for the survival difference at a given time point (in the above sense), provided the censoring mechanism can be correctly modeled, and one doubly robust method of estimation for the full survival curve. We provide R code to use these estimators for estimation and inference in the supporting information.

stat.ME

Improved small-sample inference for functions of parameters in the k-sample multinomial problem

When the target parameter for inference is a real-valued, continuous function of probabilities in the $k$-sample multinomial problem, variance estimation may be challenging. In small samples or when the function is nondifferentiable at the true parameter, methods like the nonparametric bootstrap or delta method may perform poorly. We develop an exact inference method that applies to this general situation. We prove that our proposed exact p-value correctly bounds the type I error rate and the associated confidence intervals provide at least nominal coverage; however, they are generally difficult to implement. Thus, we propose a Monte Carlo implementation to estimate the exact p-value and confidence intervals that we show to be consistent as the number of iterations grows. Our approach is general in that it applies to any real-valued continuous function of multinomial probabilities from an arbitrary number of samples and with different numbers of categories.

stat.CO

Adding covariates to bounds: What is the question?

Symbolic nonparametric bounds for partial identification of causal effects now have a long history in the causal literature. Sharp bounds, bounds that use all available information to make the range of values as narrow as possible, are often the goal. For this reason, many publications have focused on deriving sharp bounds, but the concept of sharp bounds is nuanced and can be misleading. In settings with ancillary covariates, the situation becomes more complex. We provide clear definitions for pointwise and uniform sharpness of covariate-conditional bounds, that we then use to prove some general and some specific to the IV setting results about the relationship between these two concepts. As we demonstrate, general conditions are much more difficult to determine and thus, we urge authors to be clear when including ancillary covariates in bounds via conditioning about the setting of interest and the assumptions made.

stat.ME

Inverse probability of treatment weighting with generalized linear outcome models for doubly robust estimation

There are now many options for doubly robust estimation; however, there is a concerning trend in the applied literature to believe that the combination of a propensity score and an adjusted outcome model automatically results in a doubly robust estimator and/or to misuse more complex established doubly robust estimators. A simple alternative, canonical link generalized linear models (GLM) fit via inverse probability of treatment (propensity score) weighted maximum likelihood estimation followed by standardization (the g-formula) for the average causal effect, is a doubly robust estimation method. Our aim is for the reader not just to be able to use this method, which we refer to as IPTW GLM, for doubly robust estimation, but to fully understand why it has the doubly robust property. For this reason, we define clearly, and in multiple ways, all concepts needed to understand the method and why it is doubly robust. In addition, we want to make very clear that the mere combination of propensity score weighting and an adjusted outcome model does not generally result in a doubly robust estimator. Finally, we hope to dispel the misconception that one can adjust for residual confounding remaining after propensity score weighting by adjusting in the outcome model for what remains `unbalanced' even when using doubly robust estimators. We provide R code for our simulations and real open-source data examples that can be followed step-by-step to use and hopefully understand the IPTW GLM method. We also compare to a much better-known but still simple doubly robust estimator.

stat.ME

Sharp symbolic nonparametric bounds for measures of benefit in observational and imperfect randomized studies with ordinal outcomes

The probability of benefit is a valuable and important measure of treatment effect, which has advantages over the average treatment effect. Particularly for an ordinal outcome, it has a better interpretation and can make apparent different aspects of the treatment impact. Unfortunately, this measure, and variations of it, are not identifiable even in randomized trials with perfect compliance. There is, for this reason, a long literature on nonparametric bounds for unidentifiable measures of benefit. These have primarily focused on perfect randomized trial settings and one or two specific estimands. We expand these bounds to observational settings with unmeasured confounders and imperfect randomized trials for all three estimands considered in the literature: the probability of benefit, the probability of no harm, and the relative treatment effect.

stat.ME

Accessible Computation of Tight Symbolic Bounds on Causal Effects using an Intuitive Graphical Interface

Strong untestable assumptions are almost universal in causal point estimation. In particular settings, bounds can be derived to narrow the possible range of a causal effect. Symbolic bounds apply to all settings that can be depicted using the same directed acyclic graph (DAG) and for the same effect of interest. Although the core of the methodology for deriving symbolic bounds has been previously developed, the means of implementation and computation have been lacking. Our R-package causaloptim aims to solve this usability problem by implementing the method of Sachs et al. (2022a) and providing the user with a graphical interface through shiny that allows for input in a way that most researchers with an interest in causal inference will be familiar; a DAG (via a point-and-click experience) and specifying a causal effect of interest using familiar counterfactual notation.

stat.CO

Flexible evaluation of surrogacy in Bayesian adaptive platform studies

Trial level surrogates are useful tools for improving the speed and cost effectiveness of trials, but surrogates that have not been properly evaluated can cause misleading results. The evaluation procedure is often contextual and depends on the type of trial setting. There have been many proposed methods for trial level surrogate evaluation, but none, to our knowledge, for the specific setting of Bayesian adaptive platform studies. As adaptive studies are becoming more popular, methods for surrogate evaluation using them are needed. These studies also offer a rich data resource for surrogate evaluation that would not normally be possible. However, they also offer a set of statistical issues including heterogeneity of the study population, treatments, implementation, and even potentially the quality of the surrogate. We propose the use of a hierarchical Bayesian semiparametric model for the evaluation of potential surrogates using nonparametric priors for the distribution of true effects based on Dirichlet process mixtures. The motivation for this approach is to flexibly model relationships between the treatment effect on the surrogate and the treatment effect on the outcome and also to identify potential clusters with differential surrogate value in a data-driven manner. In simulations, we find that our proposed method is superior to a simple, but fairly standard, hierarchical Bayesian method. We demonstrate how our method can be used in a simulated illustrative example (based on the ProBio trial), in which we are able to identify clusters where the surrogate is, and is not useful. We plan to apply our method to the ProBio trial, once it is completed.

stat.ME

Sharp nonparametric bounds for decomposition effects with two binary mediators

In randomized trials, once the total effect of the intervention has been estimated, it is often of interest to explore mechanistic effects through mediators along the causal pathway between the randomized treatment and the outcome. In the setting with two sequential mediators, there are a variety of decompositions of the total risk difference into mediation effects. We derive sharp and valid bounds for a number of mediation effects in the setting of two sequential mediators both with unmeasured confounding with the outcome. We provide five such bounds in the main text corresponding to two different decompositions of the total effect, as well as the controlled direct effect, with an additional thirty novel bounds provided in the supplementary materials corresponding to the terms of twenty-four four-way decompositions. We also show that, although it may seem that one can produce sharp bounds by adding or subtracting the limits of the sharp bounds for terms in a decomposition, this almost always produces valid, but not sharp bounds that can even be completely noninformative. We investigate the properties of the bounds by simulating random probability distributions under our causal model and illustrate how they are interpreted in a real data example.

stat.ME

A General Method for Deriving Tight Symbolic Bounds on Causal Effects

A causal query will commonly not be identifiable from observed data, in which case no estimator of the query can be contrived without further assumptions or measured variables, regardless of the amount or precision of the measurements of observed variables. However, it may still be possible to derive symbolic bounds on the query in terms of the distribution of observed variables. Bounds, numeric or symbolic, can often be more valuable than a statistical estimator derived under implausible assumptions. Symbolic bounds, however, provide a measure of uncertainty and information loss due to the lack of an identifiable estimand even in the absence of data. We develop and describe a general approach for computation of symbolic bounds and characterize a class of settings in which our method is guaranteed to provide tight valid bounds. This expands the known settings in which tight causal bounds are solutions to linear programs. We also prove that our method can provide valid and possibly informative symbolic bounds that are not guaranteed to be tight in a larger class of problems. We illustrate the use and interpretation of our algorithms in three examples in which we derive novel symbolic bounds.

stat.ME

A note on Horvitz-Thompson estimators for rare subgroup analysis in the presence of interference

When there is interference, a subject's outcome depends on the treatment of others and treatment effects may take on several different forms. This situation arises often, particularly in vaccine evaluation. In settings where interference is likely, two-stage cluster randomized trials have been suggested as a means of estimating some of the causal contrast of interest. Working in the finite population setting to investigate rare and unplanned subgroup analyses using some of the estimators that have been suggested in the literature, include Horvitz-Thompson, Hajek, and what might be called the natural extension of the marginal estimators suggested in Hudgens and Halloran 2008. I define the estimands of interest conditional on individual, group and both individual and group baseline variables, giving unbiased Horvitz-Thompson style estimates for each. I also provide variance estimators for several estimators. I show that the Horvitz-Thompson (HT) type estimators are always unbiased provided at least one subject within the group or population, whatever the level of interest for the estimator, is in the subgroup of interest. This is not true of the "natural" or the Hajek style estimators, which will often be undefined for rare subgroups.

math.ST

Aim for clinical utility, not just predictive accuracy

The predictions from an accurate prognostic model can be of great interest to patients and clinicians. When predictions are reported to individuals, they may decide to take action to improve their health or they may simply be comforted by the knowledge. However, if there is a clearly defined space of actions in the clinical context, a formal decision rule based on the prediction has the potential to have a much broader impact. Even if it is not the intended use of a developed prediction model, informal decision rules can often be found in practice. The use of a prediction-based decision rule should be formalized and compared to the standard of care in a randomized trial to assess its clinical utility, however, evidence is needed to motivate such a trial. We outline how observational data can be used to propose a decision rule based on a prognostic prediction model. We then propose a framework for emulating a prediction driven trial to evaluate the utility of a prediction-based decision rule in observational data. A split-sample structure can and should be used to develop the prognostic model, define the decision rule, and evaluate its clinical utility.

q-bio.QM

Ensemble Prediction of Time to Event Outcomes with Competing Risks: A Case Study of Surgical Complications in Crohn's Disease

We develop a novel algorithm to predict the occurrence of major abdominal surgery within 5 years following Crohn's disease diagnosis using a panel of 29 baseline covariates from the Swedish population registers. We model pseudo-observations based on the Aalen-Johansen estimator of the cause-specific cumulative incidence with an ensemble of modern machine learning approaches. Pseudo-observation pre-processing easily extends all existing or new machine learning procedures to right-censored event history data. We propose pseudo-observation based estimators for the area under the time varying ROC curve, for optimizing the ensemble, and the predictiveness curve, for evaluating and summarizing predictive performance.

stat.AP