Searcharxiv⌕ Search

arXiv subjects

Edwin van den Heuvel

Publications and source records attributed to Edwin van den Heuvel.

9 recordsLinked to original sources

Beyond Conditional Averages: Estimating The Individual Causal Effect Distribution

In recent years, the field of causal inference from observational data has emerged rapidly. The literature has focused on (conditional) average causal effect estimation. When (remaining) variability of individual causal effects (ICEs) is considerable, average effects may be uninformative for an individual. The fundamental problem of causal inference precludes estimating the joint distribution of potential outcomes without making assumptions. In this work, we show that the ICE distribution is identifiable under (conditional) independence of the individual effect and the potential outcome under no exposure, in addition to the common assumptions of consistency, positivity, and conditional exchangeability. Moreover, we present a family of flexible latent variable models that can be used to study individual effect modification and estimate the ICE distribution from cross-sectional data. How such latent variable models can be applied and validated in practice is illustrated in a case study on the effect of Hepatic Steatosis on a clinical precursor to heart failure. Under the assumptions presented, we estimate that 20.6% (95% Bayesian credible interval: 8.9%, 33.6%) of the population has a harmful effect greater than twice the average causal effect.

stat.ME↗

latrend: A Framework for Clustering Longitudinal Data

Clustering of longitudinal data is used to explore common trends among subjects over time for a numeric measurement of interest. Various R packages have been introduced throughout the years for identifying clusters of longitudinal patterns, summarizing the variability in trajectories between subject in terms of one or more trends. We introduce the R package "latrend" as a framework for the unified application of methods for longitudinal clustering, enabling comparisons between methods with minimal coding. The package also serves as an interface to commonly used packages for clustering longitudinal data, including "dtwclust", "flexmix", "kml", "lcmm", "mclust", "mixAK", and "mixtools". This enables researchers to easily compare different approaches, implementations, and method specifications. Furthermore, researchers can build upon the standard tools provided by the framework to quickly implement new cluster methods, enabling rapid prototyping. We demonstrate the functionality and application of the latrend package on a synthetic dataset based on the therapy adherence patterns of patients with sleep apnea.

cs.LG↗

Flexible machine learning estimation of conditional average treatment effects: a blessing and a curse

Causal inference from observational data requires untestable identification assumptions. If these assumptions apply, machine learning (ML) methods can be used to study complex forms of causal effect heterogeneity. Recently, several ML methods were developed to estimate the conditional average treatment effect (CATE). If the features at hand cannot explain all heterogeneity, the individual treatment effects (ITEs) can seriously deviate from the CATE. In this work, we demonstrate how the distributions of the ITE and the CATE can differ when a causal random forest (CRF) is applied. We extend the CRF to estimate the difference in conditional variance between treated and controls. If the ITE distribution equals the CATE distribution, this estimated difference in variance should be small. If they differ, an additional causal assumption is necessary to quantify the heterogeneity not captured by the CATE distribution. The conditional variance of the ITE can be identified when the individual effect is independent of the outcome under no treatment given the measured features. Then, in the cases where the ITE and CATE distributions differ, the extended CRF can appropriately estimate the variance of the ITE distribution while the CRF fails to do so.

stat.ME↗

Individual causal effects from observational longitudinal studies with time-varying exposures

Causal effects may vary among individuals and can even be of opposite signs. When significant effect heterogeneity exists, the population average causal effect might be uninformative for an individual. Due to the fundamental problem of causality, individual causal effects (ICEs) cannot be retrieved from cross-sectional data. However, in crossover studies, it is accepted that ICEs can be estimated under the assumptions of no carryover effects and time invariance of potential outcomes. A generic potential-outcome formulation with appropriate statistical assumptions to identify ICEs is lacking for other longitudinal data with time-varying exposures. We present a general framework for causal effect heterogeneity in which individual-specific effect modification is parameterized with a latent variable, the receptiveness factor. If the exposure varies over time, then the repeated measurements contain information on an individual's level of this receptiveness factor. Therefore, we study the conditional distribution of the ICE given all an individual's factual information. This novel conditional random variable is called the cross-world causal effect (CWCE). For known causal structures and time-varying exposures, the variability of the CWCE reduces with an increasing number of repeated measurements. The CWCE becomes identifiable from observational data under the causal assumption of cross-world similarity of individual-effect modification (i.e. there exists an exposure strategy whose effect is affected by all latent causes). We illustrate the theory with examples in which the cause-effect relations can be parameterized as generalized linear mixed assignments.

stat.ME↗

The built-in selection bias of hazard ratios formalized

It is known that the hazard ratio lacks a useful causal interpretation. Even for data from a randomized controlled trial, the hazard ratio suffers from built-in selection bias as, over time, the individuals at risk in the exposed and unexposed are no longer exchangeable. In this work, we formalize how the observed hazard ratio evolves and deviates from the causal hazard ratio of interest in the presence of heterogeneity of the hazard of unexposed individuals (frailty) and heterogeneity in effect (individual modification). For the case of effect heterogeneity, we define the causal hazard ratio. We show that the observed hazard ratio equals the ratio of expectations of the latent variables (frailty and modifier) conditionally on survival in the world with and without exposure, respectively. Examples with gamma, inverse Gaussian and compound Poisson distributed frailty, and categorical (harming, beneficial or neutral) effect modifiers are presented for illustration. This set of examples shows that an observed hazard ratio with a particular value can arise for all values of the causal hazard ratio. Therefore, the hazard ratio can not be used as a measure of the causal effect without making untestable assumptions, stressing the importance of using more appropriate estimands such as contrasts of the survival probabilities.

math.ST↗

Bias of the additive hazard model in the presence of causal effect heterogeneity

Hazard ratios are prone to selection bias, compromising their use as causal estimands. On the other hand, the hazard difference has been shown to remain unaffected by the selection of frailty factors over time. Therefore, observed hazard differences can be used as an unbiased estimator for the causal hazard differences in the absence of confounding. However, in the presence of effect (on the hazard) heterogeneity, the hazard difference is also affected by selection. In this work, we formalize how the observed hazard difference (from a randomized controlled trial) evolves by selecting favourable levels of effect modifiers in the exposed group and thus deviates from the causal hazard difference of interest. Such selection may result in a non-linear integrated hazard difference curve even when the individual causal effects are time-invariant. Therefore, a homogeneous time-varying causal additive effect on the hazard can not be distinguished from a constant but heterogeneous causal effect. We illustrate this causal issue by studying the effect of chemotherapy on the survival time of patients suffering from carcinoma of the oropharynx using data from a clinical trial. The hazard difference can thus not be used as an appropriate measure of the causal effect without making untestable assumptions.

math.ST↗

Anomaly Detection for a Large Number of Streams: A Permutation-Based Higher Criticism Approach

Anomaly detection when observing a large number of data streams is essential in a variety of applications, ranging from epidemiological studies to monitoring of complex systems. High-dimensional scenarios are usually tackled with scan-statistics and related methods, requiring stringent modeling assumptions for proper calibration. In this work we take a non-parametric stance, and propose a permutation-based variant of the higher criticism statistic not requiring knowledge of the null distribution. This results in an exact test in finite samples which is asymptotically optimal in the wide class of exponential models. We demonstrate the power loss in finite samples is minimal with respect to the oracle test. Furthermore, since the proposed statistic does not rely on asymptotic approximations it typically performs better than popular variants of higher criticism that rely on such approximations. We include recommendations such that the test can be readily applied in practice, and demonstrate its applicability in monitoring the content uniformity of an active ingredient for a batch-produced drug product.

stat.ME↗

Clustering of longitudinal data: A tutorial on a variety of approaches

During the past two decades, methods for identifying groups with different trends in longitudinal data have become of increasing interest across many areas of research. To support researchers, we summarize the guidance from the literature regarding longitudinal clustering. Moreover, we present a selection of methods for longitudinal clustering, including group-based trajectory modeling (GBTM), growth mixture modeling (GMM), and longitudinal k-means (KML). The methods are introduced at a basic level, and strengths, limitations, and model extensions are listed. Following the recent developments in data collection, attention is given to the applicability of these methods to intensive longitudinal data (ILD). We demonstrate the application of the methods on a synthetic dataset using packages available in R.

stat.ME↗

Bayesian Gaussian Copula Graphical Modeling for Dupuytren Disease

Dupuytren disease is a fibroproliferative disorder with unknown etiology that often progresses and eventually can cause permanent contractures of the affected fingers. In this paper, we provide a computationally efficient Bayesian framework to discover potential risk factors and investigate which fingers are jointly affected. Our Bayesian approach is based on Gaussian copula graphical models, which are one potential way to discover the underlying conditional independence structure of variables in multivariate mixed data. In particular, we combine the semiparametric Gaussian copula with extended rank likelihood which is appropriate to analyse multivariate mixed data with arbitrary marginal distributions. For the graph structure learning, we construct a computationally efficient search algorithm which is a trans-dimensional MCMC algorithm based on a birth-death process. In addition, to make our statistical method easily accessible to other researchers, we have implemented our method in C++ and interfaced with R software as an R package BDgraph which is available online.

stat.AP↗