SearcharxivSearch

arXiv subjects

Benjamin Glemain

Publications and source records attributed to Benjamin Glemain.

2 recordsLinked to original sources

Time Partitioning in Target Trial Emulation

In target trial emulation, time partitioning enables researchers to handle time-varying confounders and immortal time bias with appropriate methods. Based on two clinical scenarios, this study aimed to explore issues related to time partitioning and to provide guidance for trial emulation. After formalizing the research question within the framework of structural causal models, we show how a given time partitioning may be too fine or too coarse depending on the clinical context. When the partitioning is too fine, the dimensionality of the model is unnecessarily high. When the partitioning is too coarse, the resulting causal structure may hinder effect estimation. We also show that cloning-censoring-weighting may not be valid when treatment influences outcome within study periods, and we confirm this through simulations. In conclusion, we provide practical guidance for actively specifying an appropriate time partitioning in trial emulation, rather than using the available data resolution as a default.

stat.ME

Prior-Data Fitted Networks for Causal Inference: a Simulation Study with Real-World Scenarios

Prior-Data Fitted Networks (PFNs) represent a paradigm shift in tabular data prediction. We present the principles of this new paradigm and evaluate two PFNs for estimating the average treatment effect (ATE) of a binary treatment on a binary outcome, using simulated clinical scenarios based on real-world data. We assessed TabPFN combined with causal inference procedures (g-computation and inverse probability of treatment weighting), and CausalPFN, a PFN that directly provides an ATE estimate with a credible interval. Confidence intervals for the TabPFN-based methods were derived using bootstrap resampling. We found that computation times for TabPFN were prohibitive for routine causal inference, particularly because of the need for bootstrapping to yield confidence intervals. Moreover, g-computation with TabPFN produced a highly biased estimator, partially corrected by fitting separate models for each treatment group (T-learner). CausalPFN, by contrast, was computationally efficient but exhibited poor coverage of its 95% credible interval for the ATE, due to both estimation bias and inadequate uncertainty quantification. Beyond automating model specification, some PFN variants - like CausalPFN - attempt to automate causal modeling. In the settings we evaluated, CausalPFN performed poorly. However, new algorithms of this kind continue to be developed, and their application to causal inference tasks requires further investigation.

stat.AP