SearcharxivSearch

arXiv subjects

Sofia Morelli

Publications and source records attributed to Sofia Morelli.

5 recordsLinked to original sources

Representation Learning for Semiparametric Causal Mediation Analysis under No Essential Heterogeneity

We propose a two-stage estimator for structural mediation parameters that combines deep representation learning with G-estimation under the "no essential heterogeneity" (NEH) assumption. We call the method UNIT. In the first stage,TARNet estimates the heterogeneous effect of a randomized treatment on a mediator by learning a shared covariate representation across treatment arms.The resulting conditional average treatment effect (CATE) estimate provides a plug-in approximation to the heterogeneity-dependent component of the weight function entering the G-estimating equation of Zheng and Zhou (2015), which identifies the structural parameters even in the presence of unmeasured mediator-outcome confounding. We show that more accurate first-stage representation learning can yield a more informative plug-in weight and thereby improve the precision of the structural parameter estimator. In simulations with non-Gaussian covariates and nonlinear mediator effects, TARNet weights reduce the Stage-2 standard error of the mediation coefficient by a factor of $1.45$ to $1.51$ (median across replications, $n \ge 2000$) relative to the classical approach, at no cost to bias or coverage.

stat.ML

RAPSEM: Identifying Latent Mediators Without Sequential Ignorability via a Rank-Preserving Structural Equation Model

Standard structural equation models (SEMs) are often used to identify latent mediators. However, valid inference typically relies on the strong, frequently violated Sequential Ignorability assumption. We introduce the Rank-Preserving Structural Equation Model (RAPSEM), which increases robustness through G-estimation while maintaining the measurement model's integrity through a two-stage method of moments (2SMM) for factor score corrections. RAPSEM replaces the no unmeasured mediator-outcome confounding with the weaker no unobserved effect modification assumption. By leveraging treatment randomization, RAPSEM achieves identification in a manner equivalent to instrumental variable estimation through structurally emerging instruments. Specifically, identification relies on treatment-covariate interactions that influence the mediator but have no direct effect on the outcome, allowing researchers to utilize natural heterogeneity in treatment response as a testable source of identification. We provide a robustness assessment for the core identifying assumption and establish the consistency and asymptotic normality of the resulting estimator. Simulation studies demonstrate that RAPSEM remains unbiased under unobserved confounding, whereas standard SEM yields biased results. RAPSEM achieves reasonable power for sample sizes above 500, depending on the strength of the structural instruments. The method is implemented in the accompanying rapsem R package, and its practical utility is illustrated through an empirical example from educational research. The code is available at https://github.com/PsychometricsMZ/RAPSEM.

stat.ME

Dynamic Latent Class Structural Equation Modeling: A Hands-On Tutorial for Modeling Intensive Longitudinal Data

In this tutorial, we provide a hands-on guideline on how to implement complex Dynamic Latent Class Structural Equation Models (DLCSEM) in the Bayesian software JAGS. We provide building blocks starting with simple Confirmatory Factor and Time Series analysis, and then extend these blocks to Multilevel Models and Dynamic Structural Equation Models (DSEM). Subsequently, we introduce Hidden Markov Switching Models (HMSM) and demonstrate their integration with DSEM to yield DLCSEM. Leading through the tutorial is an example from clinical psychology using data on a generalized anxiety treatment that includes scales on anxiety symptoms and the Working Alliance Inventory that measures alliance between therapists and patients. Within each block, we provide an overview, specific hypotheses we want to test, the resulting model and its implementation, as well as an interpretation of the results. The aim of this tutorial is to provide a step-by-step guide for applied researchers that enables them to use this flexible DLCSEM framework for their own analyses.

stat.ME

Climate data selection for multi-decadal wind power forecasts

Reliable wind speed data is crucial for applications such as estimating local (future) wind power. Global Climate Models (GCMs) and Regional Climate Models (RCMs) provide forecasts over multi-decadal periods. However, their outputs vary substantially, and higher-resolution models come with increased computational demands. In this study, we analyze how the spatial resolution of different GCMs and RCMs affects the reliability of simulated wind speeds and wind power, using ERA5 data as a reference. We present a systematic procedure for model evaluation for wind resource assessment as a downstream task. Our results show that higher-resolution GCMs and RCMs do not necessarily preserve wind speeds more accurately. Instead, the choice of model, both for GCMs and RCMs, is more important than the resolution or GCM boundary conditions. The IPSL model preserves the wind speed distribution particularly well in Europe, producing the most accurate wind power forecasts relative to ERA5 data.

stat.AP

The impact of the spatial resolution of wind data on multi-decadal wind power forecasts in Germany

Accurate multi-decadal wind power predictions are crucial for sustainable energy transitions but are challenged by the coarse spatial resolution of global climate models (GCMs). This study examines the impact of spatial resolution on wind power forecasts by analyzing historical wind speed outputs from ten CMIP6 GCMs in Germany, using ERA5 reanalysis as a reference. Results show that the choice of GCM is the primary influence on wind speed output, with higher resolution models partly, but not consistently, improving predictions. While high-resolution models better capture extreme wind speeds, they do not systematically improve the prediction of the whole wind speed distribution. The data set MPI-ESM1-2-HR (MPI-HR) was found to represent the wind speed distribution particularly faithfully, while the MIROC6 (JAP) data set showed substantial underestimation for the German region compared to ERA5. These findings underscore the complexity of wind speed modeling for power predictions and emphasize the need for careful GCM selection and appropriate downscaling and bias correction methods.

physics.ao-ph