SearcharxivSearch

arXiv subjects

Ian Shrier

Publications and source records attributed to Ian Shrier.

10 recordsLinked to original sources

Modifying causal models to distinguish between transient and lasting causal effects

This paper considers how to classify the effects of interventions in causal models for outcomes and exposures observed over time. First, we demonstrate the limitations of the most common uses of potential outcomes and causal directed acyclic graphs for capturing all possible interventions in a time varying framework, particularly in problems where the key question concerns interventions to maintain or change equilibrium behaviour. Second, we adopt a system and state based approach rather than a measurement-based approach to identify the causal parameters. In particular, we discuss how assumptions about the system's equilibrium and the effects of interventions on that equilibrium can allow for more specific causal interpretations and clarify the goals of design and analysis. Third, we show how the ability to identify the the causal parameters of a time varying system depends on the selection of timepoints for measuring the system's states. We address this by proposing a novel version of the null effect, which is designed to distinguish between transient and lasting causal effects.

stat.ME

Recanting witness and natural direct effects: Violations of assumptions or definitions?

There have been numerous publications on the advantages and disadvantages of estimating natural (pure) effects compared to controlled effects. One of the main criticisms of natural effects is that it requires an additional assumption for identifiability, namely that the exposure does not cause a confounder of the mediator-outcome relationship. However, every analysis in every study should begin with a research question expressed in ordinary language. Researchers then develop/use mathematical expressions or estimators to best answer these ordinary language questions. When a recanting witness is present, the paper illustrates that there are no violations of assumptions. Rather, using directed acyclic graphs, the typical estimators for natural effects are simply no longer answering any meaningful question. Although some might view this as semantics, the proposed approach illustrates why the more recent methods of path-specific effects and separable effects are more valid and transparent compared to previous methods for decomposition analysis.

stat.ME

Simulation Experiments as a Causal Problem

Simulation methods are among the most ubiquitous methodological tools in statistical science. In particular, statisticians often is simulation to explore properties of statistical functionals in models for which developed statistical theory is insufficient or to assess finite sample properties of theoretical results. We show that the design of simulation experiments can be viewed from the perspective of causal intervention on a data generating mechanism. We then demonstrate the use of causal tools and frameworks in this context. Our perspective is agnostic to the particular domain of the simulation experiment which increases the potential impact of our proposed approach. In this paper, we consider two illustrative examples. First, we re-examine a predictive machine learning example from a popular textbook designed to assess the relationship between mean function complexity and the mean-squared error. Second, we discuss a traditional causal inference method problem, simulating the effect of unmeasured confounding on estimation, specifically to illustrate bias amplification. In both cases, applying causal principles and using graphical models with parameters and distributions as nodes in the spirit of influence diagrams can 1) make precise which estimand the simulation targets , 2) suggest modifications to better attain the simulation goals, and 3) provide scaffolding to discuss performance criteria for a particular simulation design.

stat.ME

Injury risk increases minimally over a large range of changes in activity level in children

Background: Limited research exists on the association between changes in physical activity levels and injury in children. Objective: To assess how well different variations of the acute:chronic workload ratio (ACWR), a measure of change in activity, predict injury in children. Methods: We conducted a prospective cohort study using data from 1670 Danish schoolchildren measured over 5.5 years (2008 to 2014). Coupled 4-week, uncoupled 4-week, and uncoupled 5-week ACWRs were calculated using activity frequency in the past week as the acute load (numerator), and average weekly activity frequency in the past 4 or 5 weeks as the chronic load (denominator). We modelled the relationship between different ACWR variations and injury using generalized linear and generalized additive models, with and without accounting for repeated measures. Results: The prognostic relationship between the ACWR and injury risk was best represented using a generalized additive mixed model for the uncoupled 5-week ACWR. It predicted an injury risk of ~3% for ACWRs between 0.8 (activity level decreased by 20%) and 1.5 (activity level increased by 50%). When activity decreased by more than 20% (ACWR< 0.8), injury risk was lower (minimum of 1.5% at ACWR=0). When activity increased by more than 50% (ACWR > 1.5), injury risk was higher (maximum of 6% at ACWR = 5). Girls were at significantly higher risk of injury than boys. Conclusion: Increases in physical activity in children are associated with much lower injury risks compared to previous results in adults.

q-bio.QM

Implementing multiple imputation for missing data in longitudinal studies when models are not feasible: A tutorial on the random hot deck approach

Objective: Researchers often use model-based multiple imputation to handle missing at random data to minimize bias while making the best use of all available data. However, there are sometimes constraints within the data that make model-based imputation difficult and may result in implausible values. In these contexts, we describe how to use random hot deck imputation to allow for plausible multiple imputation in longitudinal studies. Study Design and Setting: We illustrate random hot deck multiple imputation using The Childhood Health, Activity, and Motor Performance School Study Denmark (CHAMPS-DK), a prospective cohort study that measured weekly sports participation for 1700 Danish schoolchildren. We matched records with missing data to several observed records, generated probabilities for matched records using observed data, and sampled from these records based on the probability of each occurring. Because imputed values are generated randomly, multiple complete datasets can be created and analyzed similar to model-based multiple imputation. Conclusion: Multiple imputation using random hot deck imputation is an alternative method when model-based approaches are infeasible, specifically where there are constraints within and between covariates.

stat.ME

Causal Simulation Experiments: Lessons from Bias Amplification

Recent theoretical work in causal inference has explored an important class of variables which, when conditioned on, may further amplify existing unmeasured confounding bias (bias amplification). Despite this theoretical work, existing simulations of bias amplification in clinical settings have suggested bias amplification may not be as important in many practical cases as suggested in the theoretical literature.We resolve this tension by using tools from the semi-parametric regression literature leading to a general characterization in terms of the geometry of OLS estimators which allows us to extend current results to a larger class of DAGs, functional forms, and distributional assumptions. We further use these results to understand the limitations of current simulation approaches and to propose a new framework for performing causal simulation experiments to compare estimators. We then evaluate the challenges and benefits of extending this simulation approach to the context of a real clinical data set with a binary treatment, laying the groundwork for a principled approach to sensitivity analysis for bias amplification in the presence of unmeasured confounding.

stat.ME

The primary importance of the research question: Implications for understanding natural versus controlled direct effects and the 'cross-world independence assumption'

When developing new interventions to minimize the harmful effects of an exposure, investigators usually target the mechanisms that mediate the causal effect of the exposure on the outcome. Predicting the causal effect of these new interventions is generally done through identifying either (1) the controlled direct effect, or (2) the pure (natural) direct effect. In this opinion piece, we use the interventionist approach to discuss how these two approaches answer different questions, and the additional underlying assumptions of each compared to the other. We use a specific example for the development of a new intervention that might reduce the harmful effects of smoking on chronic obstructive pulmonary disease by removing the inhalation of harmful chemicals.

stat.ME

The acute:chronic workload ratio: challenges and prospects for improvement

Injuries occur when an athlete performs a greater amount of activity (workload) than what their body can absorb. To maximize the positive effects of training while avoiding injuries, athletes and coaches need to determine safe workload levels. The International Olympic Committee has recommended using the acute:chronic workload ratio (ACRatio) to monitor injury risk, and has provided thresholds to minimize risk. However, there are several limitations to the ACRatio which may impact the validity of current recommendations. In this review, we discuss previously published and novel challenges with the ACRatio, and possible strategies to address them. These challenges include 1) formulating the ACRatio as a proportion rather than a measure of change, 2) its use of unweighted averages to measure activity loads, 3) inapplicability of the ACRatio to sports where athletes taper their activity, 4) discretization of the ACRatio prior to model selection, 5) the establishment of the model using sparse data, 6) potential bias in the ACRatio of injured athletes, 7) unmeasured confounding, and 8) application of the ACRatio to subsequent injuries.

stat.AP

Bayesian Nonparametric Modeling of Heterogeneous Groups of Censored Data

Datasets containing large samples of time-to-event data arising from several small heterogeneous groups are commonly encountered in statistics. This presents problems as they cannot be pooled directly due to their heterogeneity or analyzed individually because of their small sample size. Bayesian nonparametric modelling approaches can be used to model such datasets given their ability to flexibly share information across groups. In this paper, we will compare three popular Bayesian nonparametric methods for modelling the survival functions of heterogeneous groups. Specifically, we will first compare the modelling accuracy of the Dirichlet process, the hierarchical Dirichlet process, and the nested Dirichlet process on simulated datasets of different sizes, where group survival curves differ in shape or in expectation. We, then, will compare the models on a real-world injury dataset.

stat.ML

A causal inference approach to network meta-analysis

While standard meta-analysis pools the results from randomized trials that compare two treatments, network meta-analysis aggregates the results of randomized trials comparing a wider variety of treatment options. However, it is unclear whether the aggregation of effect estimates across heterogeneous populations will be consistent for a meaningful parameter when not all treatments are evaluated on each population. Drawing from counterfactual theory and the causal inference framework, we define the population of interest in a network meta-analysis and define the target parameter under a series of nonparametric structural assumptions. This allows us to determine the requirements for identifiability of this parameter, enabling a description of the conditions under which network meta-analysis is appropriate and when it might mislead decision making. We then adapt several modeling strategies from the causal inference literature to obtain consistent estimation of the intervention-specific mean outcome and model-independent contrasts between treatments. Finally, we perform a reanalysis of a systematic review to compare the efficacy of antibiotics on suspected or confirmed methicillin-resistant \emph{Staphylococcus aureus} in hospitalized patients.

stat.ME