SearcharxivSearch

arXiv subjects

Wenna Xi

Publications and source records attributed to Wenna Xi.

3 recordsLinked to original sources

Clustering Informed Inverse Probability Weighting Strategies for Causal Effect Estimation in Observational Studies

Inverse probability weighting (IPW) is widely used to estimate causal effects in observational studies but depends on adequate propensity-score specification. We compare three strategies for addressing treatment assignment heterogeneity: standard IPW, clustering augmented IPW with cluster specific propensity score models, and a global propensity score model including estimated cluster membership as a covariate. Through simulations with and without latent cluster structure and under correctly specified and omitted covariate propensity score models, we evaluate bias, mean squared error (MSE), and confidence interval coverage across sample sizes of 100 to 500. Both cluster informed strategies reduced bias and MSE from omitted covariate misspecification relative to standard IPW, but neither uniformly dominated: clustering augmented IPW achieved lower MSE when latent cluster structure was present, whereas the global model generally provided lower bias and better coverage at smaller sample sizes. We also apply the methods to 966 breast cancer patients treated with carboplatin, using generalized propensity scores to estimate the dose response relationship between treatment cycles and hypersensitivity reaction risk. Standard and clustered analyses produced similar pooled estimates, while clustering additionally provided subgroup specific estimates and diagnostic profiles. Overall, cluster informed strategies may improve robustness to propensity score misspecification, with relative performance depending on subgroup structure, sample size, and inferential priorities.

stat.ME

Computing with R-INLA: Accuracy and reproducibility with implications for the analysis of COVID-19 data

The statistical methods used to analyze medical data are becoming increasingly complex. Novel statistical methods increasingly rely on simulation studies to assess their validity. Such assessments typically appear in statistical or computational journals, and the methodology is later introduced to the medical community through tutorials. This can be problematic if applied researchers use the methodologies in settings that have not been evaluated. In this paper, we explore a case study of one such method that has become popular in the analysis of coronavirus disease 2019 (COVID-19) data. The integrated nested Laplace approximations (INLA), as implemented in the R-INLA package, approximates the marginal posterior distributions of target parameters that would have been obtained from a fully Bayesian analysis. We seek to answer an important question: Does existing research on the accuracy of INLA's approximations support how researchers are currently using it to analyze COVID-19 data? We identify three limitations to work assessing INLA's accuracy: 1) inconsistent definitions of accuracy, 2) a lack of studies validating how researchers are actually using INLA, and 3) a lack of research into the reproducibility of INLA's output. We explore the practical impact of each limitation with simulation studies based on models and data used in COVID-19 research. Our results suggest existing methods of assessing the accuracy of the INLA technique may not support how COVID-19 researchers are using it. Guided in part by our results, we offer a proposed set of minimum guidelines for researchers using statistical methodologies primarily validated through simulation studies.

stat.AP

Beyond Activity Space: Detecting Communities in Ecological Networks

Emerging research suggests that the extent to which activity spaces -- the collection of an individual's routine activity locations -- overlap provides important information about the functioning of a city and its neighborhoods. To study patterns of overlapping activity spaces, we draw on the notion of an ecological network, a type of two-mode network with the two modes being individuals and the geographic locations where individuals perform routine activities. We describe a method for detecting "ecological communities" within these networks based on shared activity locations among individuals. Specifically, we identify latent activity pattern profiles, which, for each community, summarize its members' probability distribution of going to each location, and community assignment vectors, which, for each individual, summarize his/her probability distribution of belonging to each community. Using data from the Adolescent Health and Development in Context (AHDC) Study, we employ latent Dirichlet allocation (LDA) to identify activity pattern profiles and communities. We then explore differences across neighborhoods in the strength, and within-neighborhood consistency of community assignment. We hypothesize that these aspects of the neighborhood structure of ecological community membership capture meaningful dimensions of neighborhood functioning likely to co-vary with economic and racial composition. We discuss the implications of a focus on ecological communities for the conduct of "neighborhood effects" research more broadly.

stat.AP