Searcharxiv⌕ Search

arXiv subjects

Kayleigh P. Keller

Publications and source records attributed to Kayleigh P. Keller.

4 recordsLinked to original sources

Causal Inference with Multiple Misclassified Exposures: A Control Variate-Adjusted Calibration Weighting Approach

Exposure misclassification is a common issue prevalent in studies of respiratory infections in cystic fibrosis. Throat swabs are frequently used in place of expectorated or induced sputum cultures, although they have imperfect sensitivity and specificity to detect Pseudomonas aeruginosa and Staphylococcus aureus. We develop calibration weighting and control variate estimators for causal inference with multiple misclassified binary exposures and clustered observations. The calibration approach treats misclassification as a missing data problem, achieving consistency without modeling the misclassification mechanism. The control variate adjustment integrates information from error-prone observations to reduce variance while preserving the consistency of the gold-standard estimator. We show that the resulting estimator inherits double robustness from its component estimators. We also characterize a structural ceiling on efficiency gains in the bivariate setting, where joint correct classification of both exposures limits the variance reduction achievable relative to univariate applications. Simulation studies confirm the consistency and double robustness of the proposed estimators under model misspecification. We then apply these methods to a cohort of $651$ cystic fibrosis patients ages $6$--$21$. Swab-based estimates attenuate the effect of P. aeruginosa on percent predicted FEV$_1$ by approximately $79\%$ relative to sputum-based estimates ($-1.68$ vs. $-7.88$ percentage points; $95\%$ confidence interval for sputum: $-10.08$, $-5.67$). These findings suggest that relying on throat swabs may lead to under-treatment of P. aeruginosa infections. More broadly, the methods provide a framework for causal inference with multiple misclassified exposures.

stat.ME↗

Distributed Lag Interaction Model with Index Modification

Epidemiological evidence supports an association between exposure to air pollution during pregnancy and birth and child health outcomes. Typically, such associations are estimated by regressing an outcome on daily or weekly measures of exposure during pregnancy using a distributed lag model. However, these associations may be modified by multiple factors. We propose a distributed lag interaction model with index modification that allows for effect modification of a functional predictor by a weighted average of multiple modifiers. Our model allows for simultaneous estimation of modifier index weights and the exposure-time-response function via a spline cross-basis in a Bayesian hierarchical framework. Through simulations, we showed that our model out-performs competing methods when there are multiple modifiers of unknown importance. We applied our proposed method to a Colorado birth cohort to estimate the association between birth weight and air pollution modified by a neighborhood-vulnerability index and to a Mexican birth cohort to estimate the association between birthing-parent cardio-metabolic endpoints and air pollution modified by a birthing-parent lifetime stress index.

stat.ME↗

Penalized Distributed Lag Interaction Model: Air Pollution, Birth Weight and Neighborhood Vulnerability

Maternal exposure to air pollution during pregnancy has a substantial public health impact. Epidemiological evidence supports an association between maternal exposure to air pollution and low birth weight. A popular method to estimate this association while identifying windows of susceptibility is a distributed lag model (DLM), which regresses an outcome onto exposure history observed at multiple time points. However, the standard DLM framework does not allow for modification of the association between repeated measures of exposure and the outcome. We propose a distributed lag interaction model that allows modification of the exposure-time-response associations across individuals by including an interaction between a continuous modifying variable and the exposure history. Our model framework is an extension of a standard DLM that uses a cross-basis, or bi-dimensional function space, to simultaneously describe both the modification of the exposure-response relationship and the temporal structure of the exposure data. Through simulations, we showed that our model with penalization out-performs a standard DLM when the true exposure-time-response associations vary by a continuous variable. Using a Colorado, USA birth cohort, we estimated the association between birth weight and ambient fine particulate matter air pollution modified by an area-level metric of health and social adversities from Colorado EnviroScreen.

stat.ME↗

Restricted Regression in Networks

Network regression with additive node-level random effects can be problematic when the primary interest is estimating unconditional regression coefficients and some covariates are exactly or nearly in the vector space of node-level effects. We introduce the Restricted Network Regression model, that removes the collinearity between fixed and random effects in network regression by orthogonalizing the random effects against the covariates. We discuss the change in the interpretation of the regression coefficients in Restricted Network Regression and analytically characterize the effect of Restricted Network Regression on the regression coefficients for continuous response data. We show through simulation with continuous and binary response data that Restricted Network Regression mitigates, but does not alleviate, network confounding, by providing improved estimation of the regression coefficients. We apply the Restricted Network Regression model in an analysis of 2015 Eurovision Song Contest voting data and show how the choice of regression model affects inference.

stat.ME↗