Searcharxiv⌕ Search

arXiv subjects

Paul N Zivich

Publications and source records attributed to Paul N Zivich.

14 recordsLinked to original sources

Delicatessen: Automated Estimating Equations in Python

Estimating equation theory provides a unified framework for statistical modeling and inference. Applying estimating equations has historically involved extensive matrix algebra and by-hand differentiation of complex functions. Here, we introduce delicatessen, a Python library that automates those tedious calculations, lowering the barrier to adoption in quantitative biological and life science research. To highlight the utility of delicatessen for the acceleration of scientific research, we provide illustrative examples of linear regression with outliers, estimation of a dose-response curve, and standardization of the mean. Across these and other settings, delicatessen streamlines transparent and reproducible advanced statistical procedures for modern data analysis.

stat.ME↗

Estimating equations for causal survival analysis with pooled logistic regression

Background: Pooled logistic regression models are commonly applied in survival analysis. However, the standard implementation can be computationally demanding, which is further exacerbated when using the nonparametric bootstrap for inference. To ease these computational burdens, investigators often coarsen time intervals or assume a parametric functional form for time. These approaches impose restrictive assumptions, which may not always have a well-motivated substantive justification. Methods: Here, the pooled logistic regression model is re-framed using estimating equations to simplify computations and allow for inference via the empirical sandwich variance estimator, thus avoiding the more computationally demanding bootstrap. The proposed implementation is demonstrated using two examples with publicly available data. The performance of the empirical sandwich variance estimator is illustrated using a Monte Carlo simulation study. Results: As shown in the applied examples, the proposed implementation substantially reduced run-times and could be applied without needing to coarsen the data. In the simulation study, the empirical sandwich variance estimator results in nominal confidence interval coverage. Conclusions: The implementation proposed here offers a less computationally demanding alternative to the standard implementation of pooled logistic regression without needing to impose restrictive constraints on time.

stat.ME↗

Unifying Statistical and Mathematical Modeling Through a Causal Inference Lens

Within the biological, physical, and social sciences, there are two broad quantitative traditions: statistical and mathematical modeling. Both traditions have the common pursuit of advancing our scientific knowledge, but these traditions have developed largely with distinct languages and inferential frameworks. This paper uses the notion of identification from causal inference, a field originating from the statistical modeling tradition, to develop a shared language. I first review foundational identification results for statistical models and then extend these ideas to mathematical models. Central to this framework is the use of bounds, ranges of plausible numerical values, to analyze both statistical and mathematical models. I discuss the implications of this perspective for the interpretation, comparison, and integration of different modeling approaches, and illustrate the framework with a simple pharmacodynamic model for hypertension. To conclude, I describe areas where the approach taken here should be extended in the future. By formalizing connections between statistical and mathematical modeling, this work contributes to a shared framework for quantitative science. My hope is that this work will advance interactions between these two traditions.

stat.OT↗

Novel g-computation algorithms for time-varying actions with recurrent and semi-competing events

Background: A core aspect of epidemiology is determining the impacts of potential public health interventions over time. With long follow-up periods, epidemiologists may need to consider semi-competing events, in which a terminal event, like death, precludes a non-terminal event, like hypertension. Time-varying confounding poses an additional challenge when studying time-varying interventions or actions. Existing methods do not simultaneously address semi- competing events and time-varying confounding. Methods: We propose two novel g-computation algorithms for causal effects with semi- competing events and time-varying actions. To explore performance of our novel g-computation estimators, we conducted a Monte Carlo simulation study. We then applied our estimator to investigate how cigarette smoking prevention throughout young and middle adulthood might impact prevalent hypertension using data from Waves III (aged 18-26 years) - VI (aged 39-51 years) of the National Longitudinal Study of Adolescent to Adult Health. Results: Our simulations show that the novel g-computation estimators had little bias and appropriate confidence interval coverage. They outperformed existing alternative estimators across sample sizes. In the illustrative application, the novel estimator identified a small reduction in prevalence of hypertension and risk of death in midlife had all cigarette smoking been prevented across follow-up compared to the observed smoking patterns. Conclusion: As long-running cohorts progress in age, death within the study sample will become an increasing concern for studies of aging-related outcomes, life course analyses, and investigations into chronic disease development. Our novel g-computation estimators provide a simultaneous solution.

stat.ME↗

Code Sharing in Healthcare Research: A Practical Guide and Recommendations for Good Practice

As computational analysis becomes increasingly more complex in health research, transparent sharing of analytical code is vital for reproducibility and trust. This practical guide, aligned to open science practices, outlines actionable recommendations for code sharing in healthcare research. Emphasising the FAIR (Findable, Accessible, Interoperable, Reusable) principles, the authors address common barriers and provide clear guidance to help make code more robust, reusable, and scrutinised as part of the scientific record. This supports better science and more reliable evidence for computationally-driven practice and helps to adhere to new standards and guidelines of codesharing mandated by publishers and funding bodies.

cs.CY↗

Accounting for Missing Data in Public Health Research Using a Synthesis of Statistical and Mathematical Models

Introduction: Accounting for missing data by imputing or weighting conditional on covariates relies on the variable with missingness being observed at least some of the time for all unique covariate values. This requirement is referred to as positivity and positivity violations can result in bias. Here, we review a novel approach to addressing positivity violations in the context of systolic blood pressure. Methods: To illustrate the proposed approach, we estimate the mean systolic blood pressure among children and adolescents aged 2-17 years old in the United States using data from the 2017-2018 National Health and Nutrition Examination Survey (NHANES). As blood pressure was not measured for those aged 2-7, there exists a positivity violation by design. Using a recently proposed synthesis of statistical and mathematical models, we integrate external information with NHANES to address our motivating question. Results: With the synthesis model, the estimated mean systolic blood pressure was 100.5 (95% confidence interval: 99.9, 101.0), which is notably lower than either a complete-case analysis or extrapolation from a statistical model. The synthesis results were supported by a diagnostic comparing the performance of the mathematical model in the positive region. Discussion: Positivity violations pose a threat to quantitative medical research, and standard approaches to addressing nonpositivity rely on restrictive untestable assumptions. Using a synthesis model, like the one detailed here, offers a viable alternative.

stat.AP↗

Confidence Regions for Multiple Outcomes, Effect Modifiers, and Other Multiple Comparisons

In epidemiology, some have argued that multiple comparison corrections are not necessary as there is rarely interest in the universal null hypothesis. From a parameter estimation perspective, epidemiologists may still be interested in multiple parameters. In this context, standard confidence intervals are not guaranteed to provide simultaneous coverage of more than one parameter. In other words, use of confidence intervals in these cases will understate the uncertainty due to random error. To address this challenge, one can use confidence bands, an extension of confidence intervals to parameter vectors. We illustrate the use of confidence bands in three case studies: estimation of multiple causal effects, effect measure modification by a binary variable, and effect measure modification by a continuous variable. Each example uses publicly available data is accompanied by SAS, R, and Python code. The type of confidence region reported by epidemiologists should depend on whether scientific interest is in a single parameter or a set of parameters. For sets of parameters, like in cases where multiple actions or outcomes, effect measure modification, dose-response, or other functions are of interest, sup-t confidence bands are preferred due to their statistical properties, computational simplicity, and ease of presentation.

stat.ME↗

Constructing g-computation estimators: two case studies in selection bias

G-computation is a useful estimation method that can be adapted to address various biases in epidemiology. However, these adaptations may not be obvious for some complex causal structures. This challenge is an example of the much wider issue of translating a causal diagram into a novel estimation strategy. To highlight these challenges, we consider two recent cases from the selection bias literature: treatment-induced selection and co-occurrence of biases that lack a joint adjustment set. For each case study, we show how g-computation can be adapted, describe how to implement that adaptation, show some general statistical properties, and illustrate the estimator using simulation. To simplify both the theoretical study and practical application of our estimators, we express the proposed g-computation estimators as stacked estimating equations. These examples illustrate how epidemiologists can translate identification results into a g-computation estimator and study the theoretical and finite-sample properties of a novel estimator.

stat.ME↗

Empirical sandwich variance estimator for iterated conditional expectation g-computation

Iterated conditional expectation (ICE) g-computation is an estimation approach for addressing time-varying confounding for both longitudinal and time-to-event data. Unlike other g-computation implementations, ICE avoids the need to specify models for each time-varying covariate. For variance estimation, previous work has suggested the bootstrap. However, bootstrapping can be computationally intense. Here, we present ICE g-computation as a set of stacked estimating equations. Therefore, the variance for the ICE g-computation estimator can be consistently estimated using the empirical sandwich variance estimator. Performance of the variance estimator was evaluated empirically with a simulation study. The proposed approach is also demonstrated with an illustrative example on the effect of cigarette smoking on the prevalence of hypertension. In the simulation study, the empirical sandwich variance estimator appropriately estimated the variance. When comparing runtimes between the sandwich variance estimator and the bootstrap for the applied example, the sandwich estimator was substantially faster, even when bootstraps were run in parallel. The empirical sandwich variance estimator is a viable option for variance estimation with ICE g-computation.

stat.ME↗

Synthesis estimators for positivity violations with a continuous covariate

Studies intended to estimate the effect of a treatment, like randomized trials, may not be sampled from the desired target population. To correct for this discrepancy, estimates can be transported to the target population. Methods for transporting between populations are often premised on a positivity assumption, such that all relevant covariate patterns in one population are also present in the other. However, eligibility criteria, particularly in the case of trials, can result in violations of positivity when transporting to external populations. To address nonpositivity, a synthesis of statistical and mathematical models can be considered. This approach integrates multiple data sources (e.g. trials, observational, pharmacokinetic studies) to estimate treatment effects, leveraging mathematical models to handle positivity violations. This approach was previously demonstrated for positivity violations by a single binary covariate. Here, we extend the synthesis approach for positivity violations with a continuous covariate. For estimation, two novel augmented inverse probability weighting estimators are proposed. Both estimators are contrasted with other common approaches for addressing nonpositivity. Empirical performance is compared via Monte Carlo simulation. Finally, the competing approaches are illustrated with an example in the context of two-drug versus one-drug antiretroviral therapy on CD4 T cell counts among women with HIV.

stat.ME↗

Transportability without positivity: a synthesis of statistical and simulation modeling

When estimating an effect of an action with a randomized or observational study, that study is often not a random sample of the desired target population. Instead, estimates from that study can be transported to the target population. However, transportability methods generally rely on a positivity assumption, such that all relevant covariate patterns in the target population are also observed in the study sample. Strict eligibility criteria, particularly in the context of randomized trials, may lead to violations of this assumption. Two common approaches to address positivity violations are restricting the target population and restricting the relevant covariate set. As neither of these restrictions are ideal, we instead propose a synthesis of statistical and simulation models to address positivity violations. We propose corresponding g-computation and inverse probability weighting estimators. The restriction and synthesis approaches to addressing positivity violations are contrasted with a simulation experiment and an illustrative example in the context of sexually transmitted infection testing uptake. In both cases, the proposed synthesis approach accurately addressed the original research question when paired with a thoughtfully selected simulation model. Neither of the restriction approaches were able to accurately address the motivating question. As public health decisions must often be made with imperfect target population information, model synthesis is a viable approach given a combination of empirical data and external information based on the best available knowledge.

stat.ME↗

Bridged treatment comparisons: an illustrative application in HIV treatment

Comparisons of treatments, interventions, or exposures are of central interest in epidemiology, but direct comparisons are not always possible due to practical or ethical reasons. Here, we detail a fusion approach to compare treatments across studies. The motivating example entails comparing the risk of the composite outcome of death, AIDS, or greater than a 50% CD4 cell count decline in people with HIV when assigned triple versus mono antiretroviral therapy, using data from the AIDS Clinical Trial Group (ACTG) 175 (mono versus dual therapy) and ACTG 320 (dual versus triple therapy). We review a set of identification assumptions and estimate the risk difference using an inverse probability weighting estimator that leverages the shared trial arms (dual therapy). A fusion diagnostic based on comparing the shared arms is proposed that may indicate violation of the identification assumptions. Application of the data fusion estimator and diagnostic to the ACTG trials indicates triple therapy results in a reduction in risk compared to monotherapy in individuals with baseline CD4 counts between 50 and 300 cells/mm$^3$. Bridged treatment comparisons address questions that none of the constituent data sources could address alone, but valid fusion-based inference requires careful consideration of the underlying assumptions.

stat.ME↗

Positivity: Identifiability and Estimability

Positivity, the assumption that every unique combination of confounding variables that occurs in a population has a non-zero probability of an action, can be further delineated as deterministic positivity and stochastic positivity. Here, we revisit this distinction, examine its relation to nonparametric identifiability and estimability, and discuss how to address violations of positivity assumptions. Finally, we relate positivity to recent interest in machine learning, as well as the limitations of data-adaptive algorithms for causal inference. Positivity may often be overlooked, but it remains important for inference.

stat.ME↗

Machine learning for causal inference: on the use of cross-fit estimators

Modern causal inference methods allow machine learning to be used to weaken parametric modeling assumptions. However, the use of machine learning may result in complications for inference. Doubly-robust cross-fit estimators have been proposed to yield better statistical properties. We conducted a simulation study to assess the performance of several different estimators for the average causal effect (ACE). The data generating mechanisms for the simulated treatment and outcome included log-transforms, polynomial terms, and discontinuities. We compared singly-robust estimators (g-computation, inverse probability weighting) and doubly-robust estimators (augmented inverse probability weighting, targeted maximum likelihood estimation). Nuisance functions were estimated with parametric models and ensemble machine learning, separately. We further assessed doubly-robust cross-fit estimators. With correctly specified parametric models, all of the estimators were unbiased and confidence intervals achieved nominal coverage. When used with machine learning, the doubly-robust cross-fit estimators substantially outperformed all of the other estimators in terms of bias, variance, and confidence interval coverage. Due to the difficulty of properly specifying parametric models in high dimensional data, doubly-robust estimators with ensemble learning and cross-fitting may be the preferred approach for estimation of the ACE in most epidemiologic studies. However, these approaches may require larger sample sizes to avoid finite-sample issues.

stat.ME↗