SearcharxivSearch

arXiv subjects

Charles Assaad

Publications and source records attributed to Charles Assaad.

2 recordsLinked to original sources

Missing data and cluster graphs: cluster-level missingness vs variable-level missingness

Missing data is pervasive in many scientific domains such as public health, environmental science, and the social sciences. Recoverability from missing data is typically studied using fully specified variable-level missingness models despite that, in many applications, only coarse structural information is available, for instance when variables are grouped into clusters due to limited knowledge or interpretability reasons. In this paper, we investigate recoverability from such abstract representations. We introduce two classes of cluster-based missingness graphs: the m-C-DMG, which retains variable-specific missingness indicators, and the cm-C-DMG, which aggregates missingness mechanisms at the cluster level. We formalize the notion of compatibility between these abstract graphs and underlying variable-level missingness models, and study how this abstraction affects the recoverability of probabilistic and causal queries. In particular, we give graphical conditions of recovering the joint distribution as well as graphical conditions of recovering a macro causal effect. Overall, our results clarify when cluster-level missingness information is sufficient for valid inference, and when finer-grained modeling is necessary.

stat.ME

Time Partitioning in Target Trial Emulation

In target trial emulation, time partitioning enables researchers to handle time-varying confounders and immortal time bias with appropriate methods. Based on two clinical scenarios, this study aimed to explore issues related to time partitioning and to provide guidance for trial emulation. After formalizing the research question within the framework of structural causal models, we show how a given time partitioning may be too fine or too coarse depending on the clinical context. When the partitioning is too fine, the dimensionality of the model is unnecessarily high. When the partitioning is too coarse, the resulting causal structure may hinder effect estimation. We also show that cloning-censoring-weighting may not be valid when treatment influences outcome within study periods, and we confirm this through simulations. In conclusion, we provide practical guidance for actively specifying an appropriate time partitioning in trial emulation, rather than using the available data resolution as a default.

stat.ME