SearcharxivSearch

arXiv subjects

Malcolm Risk

Publications and source records attributed to Malcolm Risk.

4 recordsLinked to original sources

Studying Competing Events with Federated Cumulative Incidence Curves

Combining electronic health record (EHR) data from multiple institutions is a valuable strategy for conducting post-market safety surveillance of medical products, but privacy concerns limit sharing individual-level data. We develop a novel federated learning (FL) method for multi-site post-market safety surveillance of medical products using competing risks data. We apply this method to study immune-related adverse events (irAEs) following treatment with immune checkpoint inhibitors (ICIs) in patients with auto-immune disease (AID). We provide an algorithm for constructing non-parametric cumulative incidence curves for competing event types, which can be used to compare exposure groups (e.g. treated and untreated) with no sharing of patient-level data across institutions. We incorporate covariate adjustment via inverse propensity weighting, and informative causal comparison using the area under cumulative incidence curves, known as restricted mean time lost. We apply our method to $N=10,281$ cancer patients with no pre-existing endocrine-related AID receiving ICIs across $K=10$ sites from the OneFlorida+ network, comparing patients with a pre-existing non-endocrine AID to those with no pre-existing AID. After covariate adjustment, we found that patients with a pre-existing non-endocrine AID lost 4.8 [95% CI: 4.3,5.2] months of event-free survival time to endocrine irAEs in the first 18 months following treatment, compared to 3.2 [95% CI: 3.1,3.3] months in the group without prior AID. As patients with prior AID were initially excluded from clinical trials of ICIs, our findings provide important new information to clinicians and patients receiving or considering ICI treatment. Our proposed non-parametric federated algorithm is the first to allow investigators to use some of the most crucial non-parametric tools for conducting postmarket safety surveillance across multiple institutions.

stat.ME

Privacy-preserving causal mediation analysis using distributed electronic health record networks

Electronic health record (EHR) networks provide unprecedented opportunities to study treatment mechanisms at scale, but mediation analyses across institutions are often hindered by privacy and governance constraints that restrict sharing of patient-level data. We developed a privacy-preserving federated mediation framework that enables estimation of natural direct and indirect effects without exchanging individual-level records across participating sites. The proposed approach integrates renewable learning with counterfactual causal mediation analysis, allowing institutions to collaboratively investigate treatment mechanisms using only low-dimensional summary statistics. Both simulation studies and the real-world application demonstrated that the federated estimator closely reproduced pooled-data results while preserving patient privacy. We applied the method to 32,146 patients in the Indiana Network for Patient Care to evaluate the extent to which body mass index (BMI) mediates the effect of GLP-1 receptor agonist on glycated hemoglobin (HbA1c) reduction. The BMI-mediated pathway accounted for only a small proportion of the overall treatment effect, suggesting that most glycemic improvement occurred through mechanisms other than weight loss.

stat.AP

Time-to-Event Modeling with Pseudo-Observations in Federated Settings

In multi-center clinical research, privacy regulations often prohibit pooling individual-level records, complicating the analysis of time-to-event data. Current federated survival methods frequently require iterative communication or rely strictly on proportional hazards (PH) assumptions or require sensitive survival information. We propose a one-shot federated framework using pseudo-observations derived from a sequentially updated Kaplan-Meier estimator and fitted via a renewable generalized estimating equation. Unlike traditional methods, our approach allows flexible link functions tailored to the target estimand and accommodates non-proportional hazards. To address site-level heterogeneity, we introduce a covariate-wise debiasing procedure that shrinks noise-driven local deviations toward the global estimate while preserving genuine site-specific effects. Simulation studies demonstrate that our framework achieves inferential accuracy comparable to pooled Cox regression and the privacy-preserving One-shot Distributed Algorithm to fit a multicenter Cox proportional hazards model (ODAC) under PH assumptions, while recovering time-varying coefficient trajectories when PH is violated. Furthermore, simulations confirm that the debiasing procedure optimizes the bias-variance trade-off, adaptively balancing global stability with the preservation of genuine site-specific deviations. Applied to pediatric obesity data from the Chicago Area Patient-Centered Outcomes Research Network (CAPriCORN) network ($N=45,865$), the model produced robust estimates of time-invariant and time-varying hazard ratios, offering a flexible, privacy-preserving alternative for collaborative survival research.

stat.AP

Distributed Kaplan-Meier Analysis via the Influence Function with Application to COVID-19 and COVID-19 Vaccine Adverse Events

During the COVID-19 pandemic, regulatory decision-making was hampered by a lack of timely and high-quality data on rare outcomes. Studying rare outcomes following infection and vaccination requires conducting multi-center observational studies, where sharing individual-level data is a privacy concern. In this paper, we conduct a multi-center observational study of thromboembolic events following COVID-19 and COVID-19 vaccination without sharing individual-level data. We accomplish this by developing a novel distributed learning method for constructing Kaplan-Meier (KM) curves and inverse propensity weighted KM curves with statistical inference. We sequentially update curves site-by-site using the KM influence function, which is a measure of the direction in which an observation should shift our estimate and so can be used to incorporate new observations without access to previous data. We show in simulations that our distributed estimator is unbiased and achieves equal efficiency to the combined data estimator. Applying our method to Beaumont Health, Spectrum Health, and Michigan Medicine data, we find a much higher covariate-adjusted incidence of blood clots after SARS-CoV-2 infection (3.13%, 95% CI: [2.93, 3.35]) compared to first COVID-19 vaccine (0.08%, 95% CI: [0.08, 0.09]). This suggests that the protection vaccines provide against COVID-19-related clots outweighs the risk of vaccine-related adverse events, and shows the potential of distributed survival analysis to provide actionable evidence for time-sensitive decision making.

stat.ME