Searcharxiv⌕ Search

arXiv subjects

Kellyn F. Arnold

Publications and source records attributed to Kellyn F. Arnold.

4 recordsLinked to original sources

Lord's 'paradox' explained: the 50-year warning on the use of 'change scores' in observational data

In 1967, Frederick Lord posed a conundrum that has confused scientists for over half a century. Subsequently named Lord's 'paradox', the puzzle centres on the observation that two different approaches to estimating the effect of an exposure on the 'change' in an outcome can produce radically different results. Approach 1 involves comparing the mean 'change score' between exposure groups and Approach 2 involves comparing the follow-up outcome between exposure groups conditional on the baseline outcome. Resolving this puzzle starts with recognising the three reasons that a variable may change value: (A) 'endogenous change', which represents autocorrelation from baseline, (B) 'random change', which represents change from transient random processes, and (C) 'exogenous change', which represents all non-endogenous, non-random change and contains all change that is potentially modifiable by other baseline variables. In observational data, neither Approach 1 nor Approach 2 can reliably estimate the causal effect of an exposure on 'exogenous change' in an outcome. Approach 1 is susceptible to diluted or opposite-sign estimates whenever the exposure causes, or is caused by, the baseline outcome. Approach 2 is susceptible to inflated estimates due to measurement error in the baseline outcome and time-varying confounding bias when the baseline outcome is a mediator. The measurement error can be reduced with multiple measures of the baseline outcome, and the time-varying confounding can be reduced using g- methods. Lord's 'paradox' offers several enduring lessons for observational data science including the importance of a well-defined research question and the problems with analysing change scores in observational data.

stat.ME↗

Depicting deterministic variables within directed acyclic graphs (DAGs): An aid for identifying and interpreting causal effects involving tautological associations, compositional data, and composite variables

Deterministic variables are variables that are fully explained by one or more parent variables. They commonly arise when a variable has been algebraically constructed from one or more parent variables, as with composite variables, and in compositional data, where the 'whole' variable is determined from its 'parts'. This article introduces how deterministic variables may be depicted within directed acyclic graphs (DAGs) to help with identifying and interpreting causal effects involving tautological associations, compositional data, and composite variables. We propose a two-step approach in which all variables are initially considered, and an explicit choice is then made whether to focus on the deterministic variable(s) or the determining parents. Depicting deterministic variables within DAGs bring several benefits. It is easier to identify and avoid misinterpreting tautological associations, i.e., self-fulfilling associations between variables with shared algebraic parent variables. In compositional data, it is easier to understand the consequences of conditioning on the 'whole' variable, and correctly identify total and relative causal effects. For composite variables, it encourages greater consideration of the target estimand and greater scrutiny of the consistency and exchangeability assumptions. DAGs with deterministic variables are a useful aid for planning and interpreting analyses involving tautological associations, compositional data, and/or composite variables.

stat.ME↗

Generalised linear models for prognosis and intervention: Theory, practice, and implications for machine learning

Prediction and causal explanation are fundamentally distinct tasks of data analysis. In health applications, this difference can be understood in terms of the difference between prognosis (prediction) and prevention/treatment (causal explanation). Nevertheless, these two concepts are often conflated in practice. We use the framework of generalised linear models (GLMs) to illustrate that predictive and causal queries require distinct processes for their application and subsequent interpretation of results. In particular, we identify five primary ways in which GLMs for prediction differ from GLMs for causal inference: (1) The covariates that should be considered for inclusion in (and possibly exclusion from) the model; (2) How a suitable set of covariates to include in the model is determined; (3) Which covariates are ultimately selected, and what functional form (i.e. parameterisation) they take; (4) How the model is evaluated; and (5) How the model is interpreted. We outline some of the potential consequences of failing to acknowledge and respect these differences, and additionally consider the implications for machine learning (ML) methods. We then conclude with three recommendations which we hope will help ensure that both prediction and causal modelling are used appropriately and to greatest effect in health research.

stat.AP↗

Analyses of 'change scores' do not estimate causal effects in observational data

Background: In longitudinal data, it is common to create 'change scores' by subtracting measurements taken at baseline from those taken at follow-up, and then to analyse the resulting 'change' as the outcome variable. In observational data, this approach can produce misleading causal effect estimates. The present article uses directed acyclic graphs (DAGs) and simple simulations to provide an accessible explanation of why change scores do not estimate causal effects in observational data. Methods: Data were simulated to match three general scenarios where the variable representing measurements of the outcome at baseline was a 1) competing exposure, 2) confounder, or 3) mediator for the total causal effect of the exposure on the variable representing measurements of the outcome at follow-up. Regression coefficients were compared between change-score analyses and DAG-informed analyses. Results: Change-score analyses do not provide meaningful causal effect estimates unless the variable representing measurements of the outcome at baseline is a competing exposure, as in a randomised experiment. Where such variables (i.e. baseline measurements of the outcome) are confounders or mediators, the conclusions drawn from analyses of change scores diverge (potentially substantially) from those of DAG-informed analyses. Conclusions: Future observational studies that seek causal effect estimates should avoid analysing change scores and adopt alternative analytical strategies.

stat.ME↗