SearcharxivSearch

arXiv subjects

Danielle Braun

Publications and source records attributed to Danielle Braun.

At least 19 recordsLinked to original sources

Estimating the Effects of Heatwaves on Health: A Causal Inference Framework

The harmful relationship between heatwaves and health has been extensively documented in medical and epidemiological literature. However, most evidence is associational and cannot be interpreted causally unless strong assumptions are made. In this paper, we first make explicit the assumptions underlying the statistical methods frequently used in the heatwave literature and demonstrate when these assumptions might break down in heatwave contexts. To address these shortcomings, we propose a causal inference framework that transparently elicits causal identification assumptions. Within this new framework, we first introduce synthetic controls (SC) for estimating heatwave effects, then propose a spatially augmented Bayesian synthetic control (SA-SC) method that accounts for spatial dependence and spillovers. Empirical Monte Carlo simulations show both methods perform well, with SA-SC reducing root mean squared error and improving posterior interval coverage under spillovers and spatial dependence. Finally, we apply the proposed methods to estimate the causal effects of heatwaves on Medicare heat-related hospitalizations among 13,753,273 beneficiaries residing in Northeastern U.S. from 2000 to 2019. This causal inference framework provides spatially coherent counterfactual outcomes and robust, interpretable, and transparent causal estimates while explicitly addressing the unexamined assumptions in existing methods that pervade the heatwave effect literature.

stat.ME

A web-based user interface for Fam3PRO, a multi-gene, multi-cancer risk prediction model for families with cancer history

Purpose: Hereditary cancer risk is key to guiding screening and prevention strategies. Cancer risks can vary by individual due to the presence or absence of high- and moderate-risk pathogenic variants (PV) in cancer-associated genes, in addition to sex, age, and other risk factors. We previously developed Fam3PRO, a flexible multi-gene, multi-cancer Mendelian risk prediction model that estimates a patient's risk of carrying a PV in hereditary cancer genes and their future risk of developing several types of cancer. The Fam3PRO R package includes 22 genes with 18 associated cancers, allowing users to build customized sub-models from any gene-cancer set. However, the current R package lacks a user interface (UI), limiting its practical use in clinical settings. Therefore, we aim to develop a web-based UI for broader use of the Fam3PRO functionalities. Methods: The Fam3PRO UI (F3PI), built with R Shiny, collects and formats inputs including family health history, genetic test results, and other risk factors. Pedigree data are interactively visualized and modified via pedigreejs, while the backend Fam3PRO model takes all the inputs to generate carrier probabilities and future cancer risks, presented through an interactive UI. Results: F3PI streamlines the collection of patient and family history data, which is analyzed by the Fam3PRO models to provide personalized cancer risks for each proband across 18 cancers, as well as probabilities that a proband has a PV in up to 22 hereditary cancer genes. These results are returned to the user, within one minute on average and are available in both interactive and downloadable formats. Conclusion: We have developed F3PI, an easy-to-use, interactive web application that makes cancer and genetic risk information more accessible to providers and their patients.

stat.AP

Long-term impact of PM2.5 on mortality is exacerbated when wildfire events occur

There is extensive evidence that long-term exposure to all-source PM2.5 increases mortality. However, to date, no study has evaluated whether this effect is exacerbated in the presence of wildfire events. Here, we study 60+ million older US adults and find that wildfire events increase the harmful effects of long-term all-source PM2.5 exposure on mortality, providing a new and realistic conceptualization of wildfire health risks.

q-bio.PE

BreakLoops: A New Feature for the Multi-Gene, Multi-Cancer Family History-Based Model, Fam3Pro

Previously, we presented PanelPRO, now known as Fam3PRO, an open-source R package for multi-gene, multi-cancer risk modeling with pedigree data. The initial release could not handle pedigrees that contained cyclic structures called loops, which occur when relatives mate. Here, we present a graph-based function called breakloops that can detect and break loops in any pedigree. The core algorithm identifies the optimal set of loop breakers when individuals in a loop have exactly one parental mating, and extends to handle cases where individuals have multiple parental matings. The algorithm transforms complex pedigrees by strategically creating clones of key individuals to disrupt cycles while minimizing computational complexity. Our extensive testing demonstrates that this new feature can handle a wide variety of pedigree structures. The breakloops function is available in Fam3Pro version 2.0.0. This advancement enables Fam3Pro to assess cancer risk in a wider range of family structures, enhancing its applicability in clinical settings

stat.CO

The penetrance R package for Estimation of Age Specific Risk in Family-based Studies

Reliable tools and software for penetrance (age-specific risk among those who carry a genetic variant) estimation are critical to improving clinical decision making and risk assessment for hereditary syndromes. We introduce penetrance, an open-source R package available on CRAN, to estimate age-specific penetrance using family-history pedigree data. The package employs a Bayesian estimation approach, allowing for the incorporation of prior knowledge through the specification of priors for the parameters of the carrier distribution. It also includes options to impute missing ages during the estimation process, addressing incomplete age information which is not uncommon in pedigree datasets. Our open-source software provides a flexible and user-friendly tool for researchers to estimate penetrance in complex family-based studies, facilitating improved genetic risk assessment in hereditary syndromes.

stat.CO

Treatment Effect Heterogeneity and Importance Measures for Multivariate Continuous Treatments

Estimating the joint effect of a multivariate, continuous exposure is crucial, particularly in environmental health where interest lies in simultaneously evaluating the impact of multiple environmental pollutants on health. We develop novel methodology that addresses two key issues for estimation of treatment effects of multivariate, continuous exposures. We use nonparametric Bayesian methodology that is flexible to ensure our approach can capture a wide range of data generating processes. Additionally, we allow the effect of the exposures to be heterogeneous with respect to covariates. Treatment effect heterogeneity has not been well explored in the causal inference literature for multivariate, continuous exposures, and therefore we introduce novel estimands that summarize the nature and extent of the heterogeneity, and propose estimation procedures for new estimands related to treatment effect heterogeneity. We provide theoretical support for the proposed models in the form of posterior contraction rates and show that it works well in simulated examples both with and without heterogeneity. Our approach is motivated by a study of the health effects of simultaneous exposure to the components of PM$_{2.5}$, where we find that the negative health effects of exposure to environmental pollutants are exacerbated by low socioeconomic status, race and age.

stat.ME

Adjusting for Ascertainment Bias in Meta-Analysis of Penetrance for Cancer Risk

Multi-gene panel testing allows efficient detection of pathogenic variants in cancer susceptibility genes including moderate-risk genes such as ATM and PALB2. A growing number of studies examine the risk of breast cancer (BC) conferred by pathogenic variants of such genes. A meta-analysis combining the reported risk estimates can provide an overall age-specific risk of developing BC, i.e., penetrance for a gene. However, estimates reported by case-control studies often suffer from ascertainment bias. Currently there are no methods available to adjust for such ascertainment bias in this setting. We consider a Bayesian random-effects meta-analysis method that can synthesize different types of risk measures and extend it to incorporate studies with ascertainment bias. This is achieved by introducing a bias term in the model and assigning appropriate priors. We validate the method through a simulation study and apply it to estimate BC penetrance for carriers of pathogenic variants of ATM and PALB2 genes. Our simulations show that the proposed method results in more accurate and precise penetrance estimates compared to when no adjustment is made for ascertainment bias or when such biased studies are discarded from the analysis. The estimated overall BC risk for individuals with pathogenic variants in (1) ATM is 5.77% (3.22%-9.67%) by age 50 and 26.13% (20.31%-32.94%) by age 80; (2) PALB2 is 12.99% (6.48%-22.23%) by age 50 and 44.69% (34.40%-55.80%) by age 80. The proposed method allows for meta-analyses to include studies with ascertainment bias resulting in a larger number of studies included and thereby more robust estimates.

stat.ME

Partial identification and unmeasured confounding with multiple treatments and multiple outcomes

Estimating the health effects of multiple air pollutants is a crucial problem in public health, but one that is difficult due to unmeasured confounding bias. Motivated by this issue, we develop a framework for partial identification of causal effects in the presence of unmeasured confounding in settings with multiple treatments and multiple outcomes. Under a factor confounding assumption, we show that joint partial identification regions for multiple estimands can be more informative than considering partial identification for individual estimands one at a time. We show how assumptions related to the strength of confounding or magnitude of plausible effect sizes for one estimand can reduce the partial identification regions for other estimands. As a special case of this result, we explore how negative control assumptions reduce partial identification regions and discuss conditions under which point identification can be obtained. We develop novel computational approaches to finding partial identification regions under a variety of these assumptions. We then estimate the causal effect of PM$_{2.5}$ components on a variety of public health outcomes in the United States Medicare cohort, where we find that the detrimental effect of certain air pollutants are robust to the potential presence of unmeasured confounding bias.

stat.ME

CausalGPS: An R Package for Causal Inference With Continuous Exposures

Quantifying the causal effects of continuous exposures on outcomes of interest is critical for social, economic, health, and medical research. However, most existing software packages focus on binary exposures. We develop the CausalGPS R package that implements a collection of algorithms to provide algorithmic solutions for causal inference with continuous exposures. CausalGPS implements a causal inference workflow, with algorithms based on generalized propensity scores (GPS) as the core, extending propensity scores (the probability of a unit being exposed given pre-exposure covariates) from binary to continuous exposures. As the first step, the package implements efficient and flexible estimations of the GPS, allowing multiple user-specified modeling options. As the second step, the package provides two ways to adjust for confounding: weighting and matching, generating weighted and matched data sets, respectively. Lastly, the package provides built-in functions to fit flexible parametric, semi-parametric, or non-parametric regression models on the weighted or matched data to estimate the exposure-response function relating the outcome with the exposures. The computationally intensive tasks are implemented in C++, and efficient shared-memory parallelization is achieved by OpenMP API. This paper outlines the main components of the CausalGPS R package and demonstrates its application to assess the effect of long-term exposure to PM2.5 on educational attainment using zip code-level data from the contiguous United States from 2000-2016.

stat.CO

Bayesian Meta-Analysis of Penetrance for Cancer Risk

Multi-gene panel testing allows many cancer susceptibility genes to be tested quickly at a lower cost making such testing accessible to a broader population. Thus, more patients carrying pathogenic germline mutations in various cancer-susceptibility genes are being identified. This creates a great opportunity, as well as an urgent need, to counsel these patients about appropriate risk reducing management strategies. Counseling hinges on accurate estimates of age-specific risks of developing various cancers associated with mutations in a specific gene, i.e., penetrance estimation. We propose a meta-analysis approach based on a Bayesian hierarchical random-effects model to obtain penetrance estimates by integrating studies reporting different types of risk measures (e.g., penetrance, relative risk, odds ratio) while accounting for the associated uncertainties. After estimating posterior distributions of the parameters via a Markov chain Monte Carlo algorithm, we estimate penetrance and credible intervals. We investigate the proposed method and compare with an existing approach via simulations based on studies reporting risks for two moderate-risk breast cancer susceptibility genes, ATM and PALB2. Our proposed method is far superior in terms of coverage probability of credible intervals and mean square error of estimates. Finally, we apply our method to estimate the penetrance of breast cancer among carriers of pathogenic mutations in the ATM gene.

stat.ME

A spatial interference approach to account for mobility in air pollution studies with multivariate continuous treatments

We develop new methodology to improve our understanding of the causal effects of multivariate air pollution exposures on public health accounting for mobility. Typically, in environmental health studies, exposure to air pollution for an individual is assigned based on their residential address, though many people spend time in different regions with potentially different levels of air pollution. To account for this, we incorporate estimates of the mobility of individuals from cell phone mobility data to obtain a more accurate estimate of their air pollution exposure. We treat this as an interference problem, where individuals in one geographic region can be affected by exposures in other regions due to mobility into those areas. We propose policy-relevant estimands and derive expressions showing the extent of bias one would obtain by ignoring individuals' mobility. We additionally highlight the benefits of the proposed interference framework relative to a measurement error framework to account for mobility. Utilizing flexible Bayesian methodology we develop novel estimation strategies to estimate causal effects that account for this spatial spillover. Lastly, we use the proposed methodology to study the health effects of ambient air pollution on mortality among Medicare enrollees in the United States

stat.ME

A Bayesian Gaussian Process for Estimating a Causal Exposure Response Curve in Environmental Epidemiology

Motivated by environmental policy questions, we address the challenges of estimation, change point detection, and uncertainty quantification of a causal exposure-response function (CERF). Under a potential outcome framework, the CERF describes the relationship between a continuously varying exposure (or treatment) and its causal effect on an outcome. We propose a new Bayesian approach that relies on a Gaussian process (GP) model to estimate the CERF nonparametrically. To achieve the desired separation of design and analysis phases, we parametrize the covariance (kernel) function of the GP to mimic matching via a Generalized Propensity Score (GPS). The hyper-parameters as well as the form of the kernel function of the GP are chosen to optimize covariate balance. Our approach achieves automatic uncertainty evaluation of the CERF with high computational efficiency, and enables change point detection through inference on derivatives of the CERF. We provide theoretical results showing the correspondence between our Bayesian GP framework and traditional approaches in causal inference for estimating causal effects of a continuous exposure. We apply the methods to 520,711 ZIP-code-level observations to estimate the causal effect of long-term exposures to PM2.5, ozone, and NO2 on all-cause mortality among Medicare enrollees in the US. A computationally efficient implementation of the proposed GP models is provided in the GPCERF R package, which is available on CRAN.

stat.ME

Evaluation of Model-Based PM$_{2.5}$ Estimates for Exposure Assessment During Wildfire Smoke Episodes in the Western U.S

Investigating the health impacts of wildfire smoke requires data on people's exposure to fine particulate matter (PM$_{2.5}$) across space and time. In recent years, it has become common to use machine learning models to fill gaps in monitoring data. However, it remains unclear how well these models are able to capture spikes in PM$_{2.5}$ during and across wildfire events. Here, we evaluate the accuracy of two sets of high-coverage and high-resolution machine learning-derived PM$_{2.5}$ data sets created by Di et al. (2021) and Reid et al. (2021). In general, the Reid estimates are more accurate than the Di estimates when compared to independent validation data from mobile smoke monitors deployed by the US Forest Service. However, both models tend to severely under-predict PM$_{2.5}$ on high-pollution days. Our findings complement other recent studies calling for increased air pollution monitoring in the western US and support the inclusion of wildfire-specific monitoring observations and predictor variables in model-based estimates of PM$_{2.5}$. Lastly, we call for more rigorous error quantification of machine-learning derived exposure data sets, with special attention to extreme events.

stat.AP

Investigating Use of Low-Cost Sensors to Increase Accuracy and Equity of Real-Time Air Quality Information

Environmental Protection Agency (EPA) air quality (AQ) monitors, the gold standard for measuring air pollutants, are sparsely positioned across the US due to their costliness. Low-cost sensors (LCS) are increasingly being used by the public to fill in the gaps in AQ monitoring; however, LCS are not as accurate as EPA monitors. In this work, we investigate factors impacting the differences between an individual's true (unobserved) exposure to fine particulate matter (PM2.5) and the exposure reported by their nearest AQ instrument, which could be either an EPA monitor or an LCS. Three factors contributing to these differences are (1) distance to the nearest AQ instrument, (2) local variability in AQ, and (3) device measurement error. We examine the contributions of each component to the overall error in reported AQ using simulations based on California data. The simulations explore different combinations of hypothetical LCS placement strategies (at schools, near major roads, and in environmentally and socioeconomically marginalized census tracts) for different numbers of LCS, with varying plausible amounts of LCS device measurement error. For each scenario, we evaluate the accuracy of daily AQ information available from individuals' nearest AQ instrument with respect to absolute errors and misclassifications of the Air Quality Index, stratified by socioeconomic and demographic characteristics. We illustrate how real-time AQ reporting could be improved (or, in some cases, worsened) by using LCS, both for the population overall and for marginalized communities specifically. This work has implications for the integration of LCS into real-time AQ reporting platforms.

stat.AP

Estimating a Causal Exposure Response Function with a Continuous Error-Prone Exposure: A Study of Fine Particulate Matter and All-Cause Mortality

Numerous studies have examined the associations between long-term exposure to fine particulate matter (PM2.5) and adverse health outcomes. Recently, many of these studies have begun to employ high-resolution predicted PM2.5 concentrations, which are subject to measurement error. Previous approaches for exposure measurement error correction have either been applied in non-causal settings or have only considered a categorical exposure. Moreover, most procedures have failed to account for uncertainty induced by error correction when fitting an exposure-response function (ERF). To remedy these deficiencies, we develop a multiple imputation framework that combines regression calibration and Bayesian techniques to estimate a causal ERF. We demonstrate how the output of the measurement error correction steps can be seamlessly integrated into a Bayesian additive regression trees (BART) estimator of the causal ERF. We also demonstrate how locally-weighted smoothing of the posterior samples from BART can be used to create a more accurate ERF estimate. Our proposed approach also properly propagates the exposure measurement error uncertainty to yield accurate standard error estimates. We assess the robustness of our proposed approach in an extensive simulation study. We then apply our methodology to estimate the effects of PM2.5 on all-cause mortality among Medicare enrollees in New England from 2000-2012.

stat.ME

Assessing the causal effects of a stochastic intervention in time series data: Are heat alerts effective in preventing deaths and hospitalizations?

The methodological development of this paper is motivated by the need to address the following scientific question: does the issuance of heat alerts prevent adverse health effects? Our goal is to address this question within a causal inference framework in the context of time series data. A key challenge is that causal inference methods require the overlap assumption to hold: each unit (i.e., a day) must have a positive probability of receiving the treatment (i.e., issuing a heat alert on that day). In our motivating example, the overlap assumption is often violated: the probability of issuing a heat alert on a cooler day is zero. To overcome this challenge, we propose a stochastic intervention for time series data which is implemented via an incremental time-varying propensity score (ItvPS). The ItvPS intervention is executed by multiplying the probability of issuing a heat alert on day $t$ -- conditional on past information up to day $t$ -- by an odds ratio $δ_t$. First, we introduce a new class of causal estimands that relies on the ItvPS intervention. We provide theoretical results to show that these causal estimands can be identified and estimated under a weaker version of the overlap assumption. Second, we propose nonparametric estimators based on the ItvPS and derive an upper bound for the variances of these estimators. Third, we extend this framework to multi-site time series using a spatial meta-analysis approach. Fourth, we show that the proposed estimators perform well in terms of bias and root mean squared error via simulations. Finally, we apply our proposed approach to estimate the causal effects of increasing the probability of issuing heat alerts on each warm-season day in reducing deaths and hospitalizations among Medicare enrollees in $2,837$ U.S. counties.

stat.AP

Statistical methods for Mendelian models with multiple genes and cancers

Risk evaluation to identify individuals who are at greater risk of cancer as a result of heritable pathogenic variants is a valuable component of individualized clinical management. Using principles of Mendelian genetics, Bayesian probability theory, and variant-specific knowledge, Mendelian models derive the probability of carrying a pathogenic variant and developing cancer in the future, based on family history. Existing Mendelian models are widely employed, but are generally limited to specific genes and syndromes. However, the upsurge of multi-gene panel germline testing has spurred the discovery of many new gene-cancer associations that are not presently accounted for in these models. We have developed PanelPRO, a flexible, efficient Mendelian risk prediction framework that can incorporate an arbitrary number of genes and cancers, overcoming the computational challenges that arise because of the increased model complexity. We implement an eleven-gene, eleven-cancer model, the largest Mendelian model created thus far, based on this framework. Using simulations and a clinical cohort with germline panel testing data, we evaluate model performance, validate the reverse-compatibility of our approach with existing Mendelian models, and illustrate its usage. Our implementation is freely available for research use in the PanelPRO R package.

stat.ME

SNIP: An Adaptation of Sorted Neighborhood Methods for Deduplicating Pedigree Data

Pedigree data contain family history information that is used to analyze hereditary diseases. These clinical data sets may contain duplicate records due to the same family visiting a clinic multiple times or a clinician entering multiple versions of the family for testing purposes. Inferences drawn from the data or using them for training or validation without removing the duplicates could lead to invalid conclusions, and hence identifying the duplicates is essential. Since family structures can be complex, existing deduplication algorithms cannot be applied directly. We first motivate the importance of deduplication by examining the impact of pedigree duplicates on the training and validation of a familial risk prediction model. We then introduce an unsupervised algorithm, which we call SNIP (Sorted NeIghborhood for Pedigrees), that builds on the sorted neighborhood method to efficiently find and classify pairwise comparisons by leveraging the inherent hierarchical nature of the pedigrees. We conduct a simulation study to assess the performance of the algorithm and find parameter configurations where the algorithm is able to accurately detect the duplicates. We then apply the method to data from the Risk Service, which includes over 300,000 pedigrees at high risk of hereditary cancers, and uncover large clusters of potential duplicate families. After removing 104,520 pedigrees (33% of original data), the resulting Risk Service dataset can now be used for future analysis, training, and validation. The algorithm is available as an R package snipR available at https://github.com/bayesmendel/snipR.

stat.AP