SearcharxivSearch

arXiv subjects

Rachel C. Nethery

Publications and source records attributed to Rachel C. Nethery.

At least 19 recordsLinked to original sources

Evaluating the effects of policy interventions subject to early adoption: A case study of prescription drug monitoring programs and opioid dispensing

Policies that require organizations to use new systems, such as prescription drug monitoring programs (PDMPs), are often implemented in phases, with an initial period of voluntary access followed by mandated compliance. This allows the policy intervention to be adopted before compliance is required (early adoption), causing outcomes to change before the mandate takes effect. When early adoption is present, the no-anticipation assumption underlying synthetic control methods (SCM) is violated, leading to biased policy effect estimates. We formalize early adoption in a potential outcomes framework for staggered policy implementation and decompose the total policy effect into early adoption and mandate components. We then propose a two-stage, early adoption-aware SCM procedure that first estimates early adoption effects using an interactive fixed effects model fit to pre-mandate data and then residualizes outcomes before applying SCM variants to estimate mandate and total policy effects. Simulations, including settings with correlation between early adoption and latent factors, show reduced bias and improved uncertainty quantification relative to conventional SCM estimators. We apply the framework to state-level PDMP policies and per-capita opioid dispensing. After accounting for early adoption, estimates suggest reductions in opioid dispensing following PDMP availability and mandates; however, the estimates are imprecise and not statistically significant.

stat.ME

Debiased Machine Learning for Conformal Prediction of Counterfactual Outcomes Under Runtime Confounding

Data-driven decision making frequently relies on predicting counterfactual outcomes. In practice, researchers commonly train counterfactual prediction models on a source dataset to inform decisions on a possibly separate target population. Conformal prediction has arisen as a popular method for producing assumption-lean prediction intervals for counterfactual outcomes that would arise under different treatment decisions in the target population of interest. However, existing methods require that every confounding factor of the treatment-outcome relationship used for training on the source data is additionally measured in the target population, risking miscoverage if important confounders are unmeasured in the target population. In this paper, we introduce a computationally efficient debiased machine learning framework that allows for valid prediction intervals when only a subset of confounders is measured in the target population, a common challenge referred to as runtime confounding. Grounded in semiparametric efficiency theory, we show the resulting prediction intervals achieve desired coverage rates with faster convergence compared to standard methods. Through numerous synthetic and semi-synthetic experiments, we demonstrate the utility of our proposed method.

stat.ML

A varying-coefficient model for characterizing duration-driven heterogeneity in flood-related health impacts

Previous work revealed associations between flood exposure and adverse health outcomes during and in the aftermath of flood events. Floods are highly heterogeneous events, largely owing to vast differences in flood durations, i.e., flash-floods versus slow-moving floods. However, little to no work has incorporated exposure duration into the modeling of flood-related health impacts or has investigated duration-driven effect heterogeneity. To address this gap, we propose an exposure duration varying coefficient modeling (EDVCM) framework for estimating exposure day-specific health effects of consecutive-day environmental exposures that vary in duration. We develop the EDVCM within an area-level self-matched study design to eliminate time-invariant confounding followed by conditional Poisson regression modeling for exposure effect estimation and adjustment of time-varying confounders. Using a Bayesian framework, we introduce duration- and exposure day-specific exposure coefficients within the conditional Poisson model and assign them a two-dimensional Gaussian process prior to allow for sharing of information across both duration and exposure day. This approach enables highly-resolved insights into duration-driven effect heterogeneity while ensuring model stability through information sharing. Through simulations, we demonstrate that the EDVCM out-performs conventional approaches in terms of both effect estimation and uncertainty quantification. We apply the EDVCM to nationwide, multi-decade Medicare claims data linked with high-resolution flood exposure measures to investigate duration-driven heterogeneity in flood effects on musculoskeletal system disease hospitalizations.

stat.AP

Fair Policy Learning under Bipartite Network Interference: Learning Fair and Cost-Effective Environmental Policies

Numerous studies have shown the harmful effects of airborne pollutants on human health. Vulnerable groups and communities often bear a disproportionately larger health burden due to exposure to airborne pollutants. Thus, there is a need to design policies that effectively reduce the public health burdens while ensuring cost-effective policy interventions. Designing policies that optimally benefit the population while ensuring equity between groups under cost constraints is a challenging statistical and causal inference problem. In the context of environmental policy this is further complicated by the fact that interventions target emission sources but health impacts occur in potentially distant communities due to atmospheric pollutant transport -- a setting known as bipartite network interference (BNI). To address these issues, we propose a fair policy learning approach under BNI. Our approach allows to learn cost-effective policies under fairness constraints even accounting for complex BNI data structures. We derive asymptotic properties and demonstrate finite sample performance via Monte Carlo simulations. Finally, we apply the proposed method to a real-world dataset linking power plant scrubber installations to Medicare health records for more than 2 million individuals in the U.S. Our method determine fair scrubber allocations to reduce mortality under fairness and cost constraints.

stat.ME

Practical considerations for Gaussian Process modeling for causal inference quasi-experimental studies with panel data

Estimating causal effects in quasi-experiments with spatio-temporal panel data often requires adjusting for unmeasured confounding that varies across space and time. Gaussian Processes (GPs) offer a flexible, nonparametric modeling approach that can account for such complex dependencies through carefully chosen covariance kernels. In this paper, we provide a practical and interpretable framework for applying GPs to causal inference in panel data settings. We demonstrate how GPs generalize popular methods such as synthetic control and vertical regression, and we show that the GP posterior mean can be represented as a weighted average of observed outcomes, where the weights reflect spatial and temporal similarity. To support applied use, we explore how different kernel choices impact both estimation performance and interpretability, offering guidance for selecting between separable and nonseparable kernels. Through simulations and application to Hurricane Katrina mortality data, we illustrate how GP models can be used to estimate counterfactual outcomes and quantify treatment effects. All code and materials are made publicly available to support reproducibility and encourage adoption. Our results suggest that GPs are a promising and interpretable tool for addressing unmeasured spatio-temporal confounding in quasi-experimental studies.

stat.ME

Excess risk of heat-related hospitalization associated with temperature and PM2.5 among older adults

Background: With rising temperatures and an aging population, understanding how to prevent heat-related illness among older adults will be increasingly crucial. Despite biological plausibility, no study to date has investigated whether fine particulate matter air pollution (PM2.5) contributes to the risk of hospitalization with a diagnosis code indicating heat-related illness, referred to as heat-related hospitalization. This study aims to fill this gap by investigating the independent and combined effects of temperature and PM2.5 on heat-related hospitalization risk. Methods: We identified Medicare fee-for-service beneficiaries in the contiguous United States who experienced a heat-related hospitalization between 2008 and 2016. Using a case-crossover design and Bayesian conditional logistic regression, we characterized the associations of temperature and PM2.5 with heat-related hospitalization. We then estimated the relative excess risk due to interaction to quantify the additive interaction of simultaneous exposure to heat and PM2.5. Results: We observed 112,969 heat-related hospitalizations. Fixing PM2.5 at the case day median, the odds ratio for increasing temperature from its case day median to the 95th percentile was 1.05 (95% CI: 1.03, 1.06). Fixing temperature at the case day median, the odds ratio for increasing PM2.5 from its median to the 95th percentile was 1.01 (95% CI: 0.99, 1.04). The relative excess risk due to interaction for simultaneous median-to-95th percentile increases in temperature and PM2.5 was 0.03 (95% CI: 0.01, 0.06). Conclusions: Our study is the first to observe synergism between temperature and PM2.5 associated with the risk of heat-related hospitalization. These findings highlight the importance of considering air pollution in effective public health and clinical interventions to prevent heat-related illness.

stat.AP

A Spatiotemporal, Quasi-experimental Causal Inference Approach to Characterize the Effects of Global Plastic Waste Export and Burning on Air Quality Using Remotely Sensed Data

Open burning of plastic waste may pose a significant threat to global health by degrading air quality, but quantitative research on this problem -- crucial for policy making -- has been stunted by lack of data. Many low- and middle-income countries, where open burning is most concerning, have little to no air quality monitoring. Here, we leverage remotely sensed data products combined with spatiotemporal causal analytic techniques to evaluate the impact of large-scale plastic waste policies on air quality. Throughout, we study Indonesia before and after 2018, when China halted its import of plastic waste, resulting in diversion of this massive waste stream to other countries. We tailor cutting-edge statistical methods to this setting, estimating effects of increased plastic waste imports on fine particulate matter (PM$_{2.5}$) near waste dump sites in Indonesia as a function of proximity to ports, an induced continuous exposure. We observe strong evidence that monthly PM$_{2.5}$increased after China's ban (2018-2019) relative to expected business-as-usual (2012-2017), with increases up to 1.68 $\mu$g/m$^3$ (95% CI = [0.72, 2.48]) when exposed to medium-high port proximity. Effects were more modest for very high port proximity exposure, possibly reflecting smaller increases in dumping/burning where government oversight is greater.

stat.AP

Towards Optimal Environmental Policies: Policy Learning under Arbitrary Bipartite Network Interference

The substantial effect of air pollution on cardiovascular disease and mortality burdens is well-established. Emissions-reducing interventions on coal-fired power plants -- a major source of hazardous air pollution -- have proven to be an effective, but costly, strategy for reducing pollution-related health burdens. Targeting the power plants that achieve maximum health benefits while satisfying realistic cost constraints is challenging. The primary difficulty lies in quantifying the health benefits of intervening at particular plants. This is further complicated because interventions are applied on power plants, while health impacts occur in potentially distant communities, a setting known as bipartite network interference (BNI). In this paper, we introduce novel policy learning methods based on Q- and A-Learning to determine the optimal policy under arbitrary BNI. We derive asymptotic properties and demonstrate finite sample efficacy in simulations. We apply our novel methods to a comprehensive dataset of Medicare claims, power plant data, and pollution transport networks. Our goal is to determine the optimal strategy for installing power plant scrubbers to minimize ischemic heart disease (IHD) hospitalizations under various cost constraints. We find that annual IHD hospitalization rates could be reduced in a range from 23.37-55.30 per 10,000 person-years through optimal policies under different cost constraints.

cs.LG

Causal exposure-response curve estimation with surrogate confounders: a study of air pollution and children's health in Medicaid claims data

In this paper, we undertake a case study to estimate a causal exposure-response function (ERF) for long-term exposure to fine particulate matter (PM$_{2.5}$) and respiratory hospitalizations in socioeconomically disadvantaged children using nationwide Medicaid claims data. These data present specific challenges. First, family income-based Medicaid eligibility criteria for children differ by state, creating socioeconomically distinct populations and leading to clustered data. Second, Medicaid enrollees' socioeconomic status, a confounder and an effect modifier of the exposure-response relationships under study, is not measured. However, two surrogates are available: median household income of each enrollee's zip code and state-level Medicaid family income eligibility thresholds for children. We introduce a customized approach for causal ERF estimation called MedMatch, building on generalized propensity score (GPS) matching methods. MedMatch adapts these methods to (1) leverage the surrogate variables to account for potential confounding and/or effect modification by socioeconomic status and (2) address practical challenges presented by differing exposure distributions across clusters. We also propose a new hyperparameter selection criterion for MedMatch and traditional GPS matching methods. Through extensive simulation studies, we demonstrate the strong performance of MedMatch relative to conventional approaches in this setting. We apply MedMatch to estimate the causal ERF between PM$_{2.5}$ and respiratory hospitalization among children in Medicaid, 2000-2012. We find a positive association, with a steeper curve at lower PM$_{2.5}$ concentrations that levels off at higher concentrations.

stat.ME

Spatio-temporal quasi-experimental methods for rare disease outcomes: The impact of reformulated gasoline on childhood hematologic cancer

Although some pollutants emitted in vehicle exhaust, such as benzene, are known to cause leukemia in adults with high exposure levels, less is known about the relationship between traffic-related air pollution (TRAP) and childhood hematologic cancer. In the 1990s, the US EPA enacted the reformulated gasoline program in select areas of the US, which drastically reduced ambient TRAP in affected areas. This created an ideal quasi-experiment to study the effects of TRAP on childhood hematologic cancers. However, existing methods for quasi-experimental analyses can perform poorly when outcomes are rare and unstable, as with childhood cancer incidence. We develop Bayesian spatio-temporal matrix completion methods to conduct causal inference in quasi-experimental settings with rare outcomes. Selective information sharing across space and time enables stable estimation, and the Bayesian approach facilitates uncertainty quantification. We evaluate the methods through simulations and apply them to estimate the causal effects of TRAP on childhood leukemia and lymphoma.

stat.AP

Difference-in-Differences under Bipartite Network Interference: A Framework for Quasi-Experimental Assessment of the Effects of Environmental Policies on Health

Pollution from coal-fired power plants has been linked to substantial health and mortality burdens in the US. In recent decades, federal regulatory policies have spurred efforts to curb emissions through various actions, such as the installation of emissions control technologies on power plants. However, assessing the health impacts of these measures, particularly over longer periods of time, is complicated by several factors. First, the units that potentially receive the intervention (power plants) are disjoint from those on which outcomes are measured (communities), and second, pollution emitted from power plants disperses and affects geographically far-reaching areas. This creates a methodological challenge known as bipartite network interference (BNI). To our knowledge, no methods have been developed for conducting quasi-experimental studies with panel data in the BNI setting. In this study, motivated by the need for robust estimates of the total health impacts of power plant emissions control technologies in recent decades, we introduce a novel causal inference framework for difference-in-differences analysis under BNI with staggered treatment adoption. We explain the unique methodological challenges that arise in this setting and propose a solution via a data reconfiguration and mapping strategy. The proposed approach is advantageous because analysis is conducted at the intervention unit level, avoiding the need to arbitrarily define treatment status at the outcome unit level, but it permits interpretation of results at the more policy-relevant outcome unit level. Using this interference-aware approach, we investigate the impacts of installation of flue gas desulfurization scrubbers on coal-fired power plants on coronary heart disease hospitalizations among older Americans over the period 2003-2014, finding an overall beneficial effect in mitigating such disease outcomes.

stat.ME

Environmental Justice Implications of Power Plant Emissions Control Policies: Heterogeneous Causal Effect Estimation under Bipartite Network Interference

Emissions generators, such as coal-fired power plants, are key contributors to air pollution and thus environmental policies to reduce their emissions have been proposed. Furthermore, marginalized groups are exposed to disproportionately high levels of this pollution and have heightened susceptibility to its adverse health impacts. As a result, robust evaluations of the heterogeneous impacts of air pollution regulations are key to justifying and designing maximally protective interventions. However, such evaluations are complicated in that much of air pollution regulatory policy intervenes on large emissions generators while resulting impacts are measured in potentially distant populations. Such a scenario can be described as that of bipartite network interference (BNI). To our knowledge, no literature to date has considered estimation of heterogeneous causal effects with BNI. In this paper, we contribute to the literature in a three-fold manner. First, we propose BNI-specific estimators for subgroup-specific causal effects and design an empirical Monte Carlo simulation approach for BNI to evaluate their performance. Second, we demonstrate how these estimators can be combined with subgroup discovery approaches to identify subgroups benefiting most from air pollution policies without a priori specification. Finally, we apply the proposed methods to estimate the effects of coal-fired power plant emissions control interventions on ischemic heart disease (IHD) among 27,312,190 US Medicare beneficiaries. Though we find no statistically significant effect of the interventions in the full population, we do find significant IHD hospitalization decreases in communities with high poverty and smoking rates.

stat.ME

Optimizing Heat Alert Issuance with Reinforcement Learning

A key strategy in societal adaptation to climate change is using alert systems to prompt preventative action and reduce the adverse health impacts of extreme heat events. This paper implements and evaluates reinforcement learning (RL) as a tool to optimize the effectiveness of such systems. Our contributions are threefold. First, we introduce a new publicly available RL environment enabling the evaluation of the effectiveness of heat alert policies to reduce heat-related hospitalizations. The rewards model is trained from a comprehensive dataset of historical weather, Medicare health records, and socioeconomic/geographic features. We use scalable Bayesian techniques tailored to the low-signal effects and spatial heterogeneity present in the data. The transition model uses real historical weather patterns enriched by a data augmentation mechanism based on climate region similarity. Second, we use this environment to evaluate standard RL algorithms in the context of heat alert issuance. Our analysis shows that policy constraints are needed to improve RL's initially poor performance. Third, a post-hoc contrastive analysis provides insight into scenarios where our modified heat alert-RL policies yield significant gains/losses over the current National Weather Service alert policy in the United States.

cs.LG

A Bayesian Spatial Berkson error approach to estimate small area opioid mortality rates accounting for population-at-risk uncertainty

Monitoring small-area geographical population trends in opioid mortality has large scale implications to informing preventative resource allocation. A common approach to obtain small area estimates of opioid mortality is to use a standard disease mapping approach in which population-at-risk estimates are treated as fixed and known. Assuming fixed populations ignores the uncertainty surrounding small area population estimates, which may bias risk estimates and under-estimate their associated uncertainties. We present a Bayesian Spatial Berkson Error (BSBE) model to incorporate population-at-risk uncertainty within a disease mapping model. We compare the BSBE approach to the naive (treating denominators as fixed) using simulation studies to illustrate potential bias resulting from this assumption. We show the application of the BSBE model to obtain 2020 opioid mortality risk estimates for 159 counties in GA accounting for population-at-risk uncertainty. Utilizing our proposed approach will help to inform interventions in opioid related public health responses, policies, and resource allocation. Additionally, we provide a general framework to improve in the estimation and mapping of health indicators.

stat.ME

Severe flooding and cause-specific hospitalization in the United States

Flooding is one of the most disruptive and costliest climate-related disasters and presents an escalating threat to population health due to climate change and urbanization patterns. Previous studies have investigated the consequences of flood exposures on only a handful of health outcomes and focus on a single flood event or affected region. To address this gap, we conducted a nationwide, multi-decade analysis of the impacts of severe floods on a wide range of health outcomes in the United States by linking a novel satellite-based high-resolution flood exposure database with Medicare cause-specific hospitalization records over the period 2000- 2016. Using a self-matched study design with a distributed lag model, we examined how cause-specific hospitalization rates deviate from expected rates during and up to four weeks after severe flood exposure. Our results revealed that risk of hospitalization was consistently elevated during and for at least four weeks following severe flood exposure for nervous system diseases (3.5 %; 95 % confidence interval [CI]: 0.6 %, 6.4 %), skin and subcutaneous tissue diseases (3.4 %; 95 % CI: 0.3 %, 6.7 %), and injury and poisoning (1.5 %; 95 % CI: -0.07 %, 3.2 %). Increases in hospitalization rate for these causes, musculoskeletal system diseases, and mental health-related impacts varied based on proportion of Black residents in each ZIP Code. Our findings demonstrate the need for targeted preparedness strategies for hospital personnel before, during, and after severe flooding.

stat.AP

Impacts of Census Differential Privacy for Small-Area Disease Mapping to Monitor Health Inequities

The US Census Bureau will implement a new privacy-preserving disclosure avoidance system (DAS), which includes application of differential privacy, on the public-release 2020 census data. There are concerns that the DAS may bias small-area and demographically-stratified population counts, which play a critical role in public health research and policy, serving as denominators in estimation of disease/mortality rates. Employing three DAS demonstration products, we quantify errors attributable to reliance on DAS-protected denominators in standard small-area disease mapping models for characterizing health inequities. We conduct simulation studies and real data analyses of inequities in premature mortality at the census tract level in Massachusetts. Results show that overall patterns of inequity by racialized group and economic deprivation level are not compromised by the DAS. While early versions of DAS induce errors in mortality rate estimation that are larger for Black than for non-Hispanic white populations, this issue is ameliorated in newer DAS versions.

stat.AP

Evaluation of Model-Based PM$_{2.5}$ Estimates for Exposure Assessment During Wildfire Smoke Episodes in the Western U.S

Investigating the health impacts of wildfire smoke requires data on people's exposure to fine particulate matter (PM$_{2.5}$) across space and time. In recent years, it has become common to use machine learning models to fill gaps in monitoring data. However, it remains unclear how well these models are able to capture spikes in PM$_{2.5}$ during and across wildfire events. Here, we evaluate the accuracy of two sets of high-coverage and high-resolution machine learning-derived PM$_{2.5}$ data sets created by Di et al. (2021) and Reid et al. (2021). In general, the Reid estimates are more accurate than the Di estimates when compared to independent validation data from mobile smoke monitors deployed by the US Forest Service. However, both models tend to severely under-predict PM$_{2.5}$ on high-pollution days. Our findings complement other recent studies calling for increased air pollution monitoring in the western US and support the inclusion of wildfire-specific monitoring observations and predictor variables in model-based estimates of PM$_{2.5}$. Lastly, we call for more rigorous error quantification of machine-learning derived exposure data sets, with special attention to extreme events.

stat.AP

Investigating Use of Low-Cost Sensors to Increase Accuracy and Equity of Real-Time Air Quality Information

Environmental Protection Agency (EPA) air quality (AQ) monitors, the gold standard for measuring air pollutants, are sparsely positioned across the US due to their costliness. Low-cost sensors (LCS) are increasingly being used by the public to fill in the gaps in AQ monitoring; however, LCS are not as accurate as EPA monitors. In this work, we investigate factors impacting the differences between an individual's true (unobserved) exposure to fine particulate matter (PM2.5) and the exposure reported by their nearest AQ instrument, which could be either an EPA monitor or an LCS. Three factors contributing to these differences are (1) distance to the nearest AQ instrument, (2) local variability in AQ, and (3) device measurement error. We examine the contributions of each component to the overall error in reported AQ using simulations based on California data. The simulations explore different combinations of hypothetical LCS placement strategies (at schools, near major roads, and in environmentally and socioeconomically marginalized census tracts) for different numbers of LCS, with varying plausible amounts of LCS device measurement error. For each scenario, we evaluate the accuracy of daily AQ information available from individuals' nearest AQ instrument with respect to absolute errors and misclassifications of the Air Quality Index, stratified by socioeconomic and demographic characteristics. We illustrate how real-time AQ reporting could be improved (or, in some cases, worsened) by using LCS, both for the population overall and for marginalized communities specifically. This work has implications for the integration of LCS into real-time AQ reporting platforms.

stat.AP