SearcharxivSearch

arXiv subjects

Erica Tavazzi

Publications and source records attributed to Erica Tavazzi.

3 recordsLinked to original sources

Treatment persistence drives estimator performance in longitudinal causal inference based on observational data: A simulation study

Longitudinal clinical data are increasingly available, offering opportunities to study treatment effects over time but also raising challenges related to time-varying confounding and evolving treatment decisions. We investigate how longitudinal treatment dynamics affect causal effect estimation when baseline and longitudinal methods target different causal estimands. Using a structural causal model (SCM), we simulate data with time-varying confounding, binary treatment, and an absorbing binary outcome under 9 scenarios combining functional complexity and treatment persistence. We compare three baseline and two longitudinal estimators against Monte Carlo ground-truth risk ratios (RRs) under sustained treatment regimes. Results show that treatment persistence is the main driver of estimator behaviour. High persistence reduces the discrepancy between baseline and sustained-regime estimands, making baseline estimators closer to the sustained-regime ground truth (average relative deviation of baseline IPTW/TMLE decreasing from 61% under low persistence to 10% under high persistence), whereas low persistence induces practical positivity challenges and increases the variability of longitudinal estimators, with empirical 95% interval widths increasing from 0.22 to 0.46 for longitudinal IPTW and from 0.20 to 0.36 for LTMLE when moving from high to low persistence. These findings emphasise that estimator performance should be interpreted jointly with the target intervention and the treatment process generating the observed data.

stat.ME

Exploring the Impact of Environmental Pollutants on Multiple Sclerosis Progression

Multiple Sclerosis (MS) is a chronic autoimmune and inflammatory neurological disorder characterised by episodes of symptom exacerbation, known as relapses. In this study, we investigate the role of environmental factors in relapse occurrence among MS patients, using data from the H2020 BRAINTEASER project. We employed predictive models, including Random Forest (RF) and Logistic Regression (LR), with varying sets of input features to predict the occurrence of relapses based on clinical and pollutant data collected over a week. The RF yielded the best result, with an AUC-ROC score of 0.713. Environmental variables, such as precipitation, NO2, PM2.5, humidity, and temperature, were found to be relevant to the prediction.

cs.LG

Comparing Propensity Score-Based Methods in Estimating the Treatment Effects: A Simulation Study

In observational studies, the recorded treatment assignment is not purely random, but it is influenced by external factors such as patient characteristics, reimbursement policies, and existing guidelines. Therefore, the treatment effect can be estimated only after accounting for confounding factors. Propensity score (PS) methods are a family of methods that is widely used for this purpose. Although they are all based on the estimation of the a posteriori probability of treatment assignment given patient covariates, they estimate the treatment effect from different statistical points of view and are, thus, relatively hard to compare. In this work, we propose a simulation experiment in which a hypothetical cohort of subjects is simulated in seven scenarios of increasing complexity of the associations between covariates and treatment, but where the two main definitions of treatment effect (average treatment effect, ATE, and average effect of the treatment on the treated, ATT) coincide. Our purpose is to compare the performance of a wide array of PS-based methods (matching, stratification, and inverse probability weighting) in estimating the treatment effect and their robustness in different scenarios. We find that inverse probability weighting provides estimates of the treatment effect that are closer to the expected value by weighting all subjects of the starting population. Conversely, matching and stratification ensure that the subpopulation that generated the final estimate is made up of real instances drawn from the starting population, and, thus, provide a higher degree of control on the validity domain of the estimates.

stat.ME