SearcharxivSearch

arXiv subjects

Robert W. Platt

Publications and source records attributed to Robert W. Platt.

14 recordsLinked to original sources

Early Pregnancy Treatment Decisions: Designing Perinatal Pharmacoepidemiology Studies using Real-World Data

Research on the use of medications during pregnancy has two primary goals: to detect signals that medications may be harmful to a pregnant individual or fetus, and to support better treatment of pregnant people who require pharmacotherapy. Target trial emulation has been proposed as an approach to estimate the effects of interventions in real world data, with recent extensions to the pregnancy setting. This approach focuses on aligning eligibility and treatment initiation with start of follow up, which is particularly desirable given methodological challenges specific to pregnancy, such as right and left censoring and truncation, competing events, differences in gestational length, and varying etiologically susceptible periods. While previous work on target trial emulation in pregnancy has focused on initiation versus non-initiation of point treatments such as vaccines or antibiotics, this paper focuses on research questions regarding changes to pregestational treatment regimes, and aims to highlight opportunities and approaches to designing studies that align with relevant time points during early pregnancy at which treatment decisions occur in clinical practice. Using the example of treatment for type 2 diabetes mellitus, we review methods for identifying pregnancy episodes in routinely collected healthcare data, introduce possible time zero candidates, and discuss analytic approaches that minimize potential bias due to selection and immortal person time.

stat.AP

Valid post-selection inference for penalized G-estimation

Understanding treatment effect heterogeneity is important for decision making in medical and clinical practices, or handling various engineering and marketing challenges. When dealing with high-dimensional covariates or when the effect modifiers are not predefined and need to be discovered, data-adaptive selection approaches become essential. However, with data-driven model selection, the quantification of statistical uncertainty is complicated by post-selection inference due to difficulties in approximating the sampling distribution of the target estimator. Data-driven model selection tends to favor models with strong effect modifiers with an associated cost of inflated type I errors. Although several frameworks and methods for valid statistical inference have been proposed for ordinary least squares regression following data-driven model selection, fewer options exist for valid inference for effect modifier discovery in causal modeling contexts. In this article, we extend two different methods to develop valid inference for penalized G-estimation that investigates effect modification of proximal treatment effects within the structural nested mean model framework. We show the asymptotic validity of the proposed methods. Using extensive simulation studies, we evaluate and compare the finite sample performance of the proposed methods and the naive inference based on a sandwich variance estimator. Our work is motivated by the study of hemodiafiltration for treating patients with end-stage renal disease at the Centre Hospitalier de l'Université de Montréal. We apply these methods to draw inference about the effect heterogeneity of dialysis facility on the repeated session-specific hemodiafiltration outcomes.

stat.ME

Is checking for sequential positivity violations getting you down? Try sPoRT!

Background: Sequential positivity is often a necessary assumption for drawing causal inferences, such as through marginal structural modeling. Unfortunately, verification of this assumption can be challenging because it usually relies on multiple parametric propensity score models, unlikely all correctly specified. Therefore, we propose a new algorithm, called sequential Positivity Regression Tree (sPoRT), to overcome this issue and identify the subgroups found to be violating this assumption, allowing for insights about the nature of the violations and potential solutions. Methods: We present different versions of sPoRT based on either stratifying or pooling over time under static or dynamic treatment strategies. This methodological development was motivated by a real-life application of the impact of the timing of initiation of HIV treatment with and without smoothing over time, which we also use to demonstrate the method. Results: The illustration of sPoRT demonstrates its easy use and the interpretability of the results for applied epidemiologists. Furthermore, an R notebook showing how to use sPoRT in practice is available at github.com/ArthurChatton/sPoRT-notebook. Conclusions: The sPoRT algorithm provides interpretable subgroups violating the sequential positivity violation, allowing patterns and trends in the confounders to be easily identified. We finally provided practical implications and recommendations when positivity violations are identified.

stat.ME

Penalized G-estimation for effect modifier selection in a structural nested mean model for repeated outcomes

Effect modification occurs when the impact of the treatment on an outcome varies based on the levels of other covariates known as effect modifiers. Modeling these effect differences is important for etiological goals and for purposes of optimizing treatment. Structural nested mean models (SNMMs) are useful causal models for estimating the potentially heterogeneous effect of a time-varying exposure on the mean of an outcome in the presence of time-varying confounding. A data-adaptive selection approach is necessary if the effect modifiers are unknown a priori and need to be identified. Although variable selection techniques are available for estimating the conditional average treatment effects using marginal structural models or for developing optimal dynamic treatment regimens, all of these methods consider a single end-of-follow-up outcome. In the context of an SNMM for repeated outcomes, we propose a doubly robust penalized G-estimator for the causal effect of a time-varying exposure with a simultaneous selection of effect modifiers and prove the oracle property of our estimator. We conduct a simulation study for the evaluation of its performance in finite samples and verification of its double-robustness property. Our work is motivated by the study of hemodiafiltration for treating patients with end-stage renal disease at the Centre Hospitalier de l'Université de Montréal. We apply the proposed method to investigate the effect heterogeneity of dialysis facility on the repeated session-specific hemodiafiltration outcomes.

stat.ME

What if we had built a prediction model with a survival super learner instead of a Cox model 10 years ago?

Objective: This study sought to compare the drop in predictive performance over time according to the modeling approach (regression versus machine learning) used to build a kidney transplant failure prediction model with a time-to-event outcome. Study Design and Setting: The Kidney Transplant Failure Score (KTFS) was used as a benchmark. We reused the data from which it was developed (DIVAT cohort, n=2,169) to build another prediction algorithm using a survival super learner combining (semi-)parametric and non-parametric methods. Performance in DIVAT was estimated for the two prediction models using internal validation. Then, the drop in predictive performance was evaluated in the same geographical population approximately ten years later (EKiTE cohort, n=2,329). Results: In DIVAT, the super learner achieved better discrimination than the KTFS, with a tAUROC of 0.83 (0.79-0.87) compared to 0.76 (0.70-0.82). While the discrimination remained stable for the KTFS, it was not the case for the super learner, with a drop to 0.80 (0.76-0.83). Regarding calibration, the survival SL overestimated graft survival at development, while the KTFS underestimated graft survival ten years later. Brier score values were similar regardless of the approach and the timing. Conclusion: The more flexible SL provided superior discrimination on the population used to fit it compared to a Cox model and similar discrimination when applied to a future dataset of the same population. Both methods are subject to calibration drift over time. However, weak calibration on the population used to develop the prediction model was correct only for the Cox model, and recalibration should be considered in the future to correct the calibration drift.

stat.ME

Personalised dynamic super learning: an application in predicting hemodiafiltration convection volumes

Obtaining continuously updated predictions is a major challenge for personalised medicine. Leveraging combinations of parametric regressions and machine learning approaches, the personalised online super learner (POSL) can achieve such dynamic and personalised predictions. We adapt POSL to predict a repeated continuous outcome dynamically and propose a new way to validate such personalised or dynamic prediction models. We illustrate its performance by predicting the convection volume of patients undergoing hemodiafiltration. POSL outperformed its candidate learners with respect to median absolute error, calibration-in-the-large, discrimination, and net benefit. We finally discuss the choices and challenges underlying the use of POSL.

stat.ME

Improving subgroup analysis using methods to extend inferences to specific target populations

Subgroup analyses are common in epidemiologic and clinical research. Unfortunately, restriction to subgroup members to test for heterogeneity can yield imprecise effect estimates. If the true effect differs between members and non-members due to different distributions of other measured effect measure modifiers (EMMs), leveraging data from non-members can improve the precision of subgroup effect estimates. We obtained data from the PRIME RCT of panitumumab in patients with metastatic colon and rectal cancer from Project Datasphere(TM) to demonstrate this method. We weighted non-Hispanic White patients to resemble Hispanic patients in measured potential EMMs (e.g., age, KRAS distribution, sex), combined Hispanic and weighted non-Hispanic White patients in one data set, and estimated 1-year differences in progression-free survival (PFS). We obtained percentile-based 95% confidence limits for this 1-year difference in PFS from 2,000 bootstraps. To show when the method is less helpful, we also reweighted male patients to resemble female patients and mutant-type KRAS (no treatment benefit) patients to resemble wild-type KRAS (treatment benefit) patients. The PRIME RCT included 795 non-Hispanic White and 42 Hispanic patients with complete data on EMMs. While the Hispanic-only analysis estimated a one-year PFS change of -17% (95% C.I. -45%, 8.8%) with panitumumab, the combined weighted estimate was more precise (-8.7%, 95% CI -22%, 5.3%) while differing from the full population estimate (1.0%, 95% CI: -5.9%, 7.5%). When targeting wild-type KRAS patients the combined weighted estimate incorrectly suggested no benefit (one-year PFS change: 0.9%, 95% CI: -6.0%, 7.2%). Methods to extend inferences from study populations to specific targets can improve the precision of estimates of subgroup effect estimates when their assumptions are met. Violations of those assumptions can lead to bias, however.

stat.AP

The Complex Estimand of Clone-Censor-Weighting When Studying Treatment Initiation Windows

Clone-censor-weighting (CCW) is an analytic method for studying treatment regimens that are indistinguishable from one another at baseline without relying on landmark dates or creating immortal person time. One particularly interesting CCW application is estimating outcomes when starting treatment within specific time windows in observational data (e.g., starting a treatment within 30 days of hospitalization). In such cases, CCW estimates something fairly complex. We show how using CCW to study a regimen such as "start treatment prior to day 30" estimates the potential outcome of a hypothetical intervention where A) prior to day 30, everyone follows the treatment start distribution of the study population and B) everyone who has not initiated by day 30 initiates on day 30. As a result, the distribution of treatment initiation timings provides essential context for the results of CCW studies. We also show that if the exposure effect varies over time, ignoring exposure history when estimating inverse probability of censoring weights (IPCW) estimates the risk under an impossible intervention and can create selection bias. Finally, we examine some simplifying assumptions that can make this complex treatment effect more interpretable and allow everyone to contribute to IPCW.

stat.ME

Integrating complex selection rules into the latent overlapping group Lasso for constructing coherent prediction models

The construction of coherent prediction models holds great importance in medical research as such models enable health researchers to gain deeper insights into disease epidemiology and clinicians to identify patients at higher risk of adverse outcomes. One commonly employed approach to developing prediction models is variable selection through penalized regression techniques. Integrating natural variable structures into this process not only enhances model interpretability but can also %increase the likelihood of recovering the true underlying model and boost prediction accuracy. However, a challenge lies in determining how to effectively integrate potentially complex selection dependencies into the penalized regression. In this work, we demonstrate how to represent selection dependencies mathematically, provide algorithms for deriving the complete set of potential models, and offer a structured approach for integrating complex rules into variable selection through the latent overlapping group Lasso. To illustrate our methodology, we applied these techniques to construct a coherent prediction model for major bleeding in hypertensive patients recently hospitalized for atrial fibrillation and subsequently prescribed oral anticoagulants. In this application, we account for a proxy of anticoagulant adherence and its interaction with dosage and the type of oral anticoagulants in addition to drug-drug interactions.

stat.AP

A general framework for formulating structured variable selection

In variable selection, a selection rule that prescribes the permissible sets of selected variables (called a "selection dictionary") is desirable due to the inherent structural constraints among the candidate variables. Such selection rules can be complex in real-world data analyses, and failing to incorporate such restrictions could not only compromise the interpretability of the model but also lead to decreased prediction accuracy. However, no general framework has been proposed to formalize selection rules and their applications, which poses a significant challenge for practitioners seeking to integrate these rules into their analyses. In this work, we establish a framework for structured variable selection that can incorporate universal structural constraints. We develop a mathematical language for constructing arbitrary selection rules, where the selection dictionary is formally defined. We demonstrate that all selection rules can be expressed as combinations of operations on constructs, facilitating the identification of the corresponding selection dictionary. Once this selection dictionary is derived, practitioners can apply their own user-defined criteria to select the optimal model. Additionally, our framework enhances existing penalized regression methods for variable selection by providing guidance on how to appropriately group variables to achieve the desired selection rule. Furthermore, our innovative framework opens the door to establishing new l0 norm-based penalized regression techniques that can be tailored to respect arbitrary selection rules, thereby expanding the possibilities for more robust and tailored model development.

stat.ME

Structured Learning in Time-dependent Cox Models

Cox models with time-dependent coefficients and covariates are widely used in survival analysis. In high-dimensional settings, sparse regularization techniques are employed for variable selection, but existing methods for time-dependent Cox models lack flexibility in enforcing specific sparsity patterns (i.e., covariate structures). We propose a flexible framework for variable selection in time-dependent Cox models, accommodating complex selection rules. Our method can adapt to arbitrary grouping structures, including interaction selection, temporal, spatial, tree, and directed acyclic graph structures. It achieves accurate estimation with low false alarm rates. We develop the sox package, implementing a network flow algorithm for efficiently solving models with complex covariate structures. sox offers a user-friendly interface for specifying grouping structures and delivers fast computation. Through examples, including a case study on identifying predictors of time to all-cause death in atrial fibrillation patients, we demonstrate the practical application of our method with specific selection rules.

stat.ME

Evaluating hybrid controls methodology in early-phase oncology trials: a simulation study based on the MORPHEUS-UC trial

Phase Ib/II oncology trials, despite their small sample sizes, aim to provide information for optimal internal company decision-making concerning novel drug development. Hybrid controls (a combination of the current control arm and controls from one or more sources of historical trial data [HTD]) can be used to increase the statistical precision. Here we assess combining two sources of Roche HTD to construct a hybrid control in targeted therapy for decision-making via an extensive simulation study. Our simulations are based on the real data of one of the experimental arms and the control arm of the MORPHEUS-UC Phase Ib/II study and two Roche HTD for atezolizumab monotherapy. We consider potential complications such as model misspecification, unmeasured confounding, different sample sizes of current treatment groups, and heterogeneity among the three trials. We evaluate two frequentist methods (with both Cox and Weibull accelerated failure time [AFT] models) and three different priors in Bayesian dynamic borrowing (with a Weibull AFT model), and modifications within each of those, when estimating the effect of treatment on survival outcomes and measures of effect such as marginal hazard ratios. We assess the performance of these methods in different settings and potential of generalizations to supplement decisions in early-phase oncology trials. The results show that the proposed joint frequentist methods and noninformative priors within Bayesian dynamic borrowing with no adjustment on covariates are preferred, especially when treatment effects across the three trials are heterogeneous. For generalization of hybrid control methods in such settings we recommend more simulation studies.

stat.ME

Estimation of the marginal effect of antidepressants on body mass index under confounding and endogenous covariate-driven monitoring times

In studying the marginal effect of antidepressants on body mass index using electronic health records data, we face several challenges. Patients' characteristics can affect the exposure (confounding) as well as the timing of routine visits (measurement process), and those characteristics may be altered following a visit which can create dependencies between the monitoring and body mass index when viewed as a stochastic or random processes in time. This may result in a form of selection bias that distorts the estimation of the marginal effect of the antidepressant. Inverse intensity of visit weights have been proposed to adjust for these imbalances, however no approaches have addressed complex settings where the covariate and the monitoring processes affect each other in time so as to induce endogeneity, a situation likely to occur in electronic health records. We review how selection bias due to outcome-dependent follow-up times may arise and propose a new cumulated weight that models a complete monitoring path so as to address the above-mentioned challenges and produce a reliable estimate of the impact of antidepressants on body mass index. More specifically, we do so using data from the Clinical Practice Research Datalink in the United Kingdom, comparing the marginal effect of two commonly used antidepressants, citalopram and fluoxetine, on body mass index. The results are compared to those obtained with simpler methods that do not account for the extent of the dependence due to an endogenous covariate process.

stat.AP

Effect of breastfeeding on gastrointestinal infection in infants: A targeted maximum likelihood approach for clustered longitudinal data

The PROmotion of Breastfeeding Intervention Trial (PROBIT) cluster-randomized a program encouraging breastfeeding to new mothers in hospital centers. The original studies indicated that this intervention successfully increased duration of breastfeeding and lowered rates of gastrointestinal tract infections in newborns. Additional scientific and popular interest lies in determining the causal effect of longer breastfeeding on gastrointestinal infection. In this study, we estimate the expected infection count under various lengths of breastfeeding in order to estimate the effect of breastfeeding duration on infection. Due to the presence of baseline and time-dependent confounding, specialized "causal" estimation methods are required. We demonstrate the double-robust method of Targeted Maximum Likelihood Estimation (TMLE) in the context of this application and review some related methods and the adjustments required to account for clustering. We compare TMLE (implemented both parametrically and using a data-adaptive algorithm) to other causal methods for this example. In addition, we conduct a simulation study to determine (1) the effectiveness of controlling for clustering indicators when cluster-specific confounders are unmeasured and (2) the importance of using data-adaptive TMLE.

stat.AP