SearcharxivSearch

arXiv subjects

Margaret Gamalo

Publications and source records attributed to Margaret Gamalo.

6 recordsLinked to original sources

Transporting Randomized Trial Effects to Real-World Populations via Riesz-Calibrated Optimal Transport

Randomized trials support causal inference, but differences between trial and target populations can limit the transportability of treatment effects to real-world settings. Many existing approaches model the propensity of trial participation and can therefore be sensitive to model misspecification and weak overlap of the covariate distributions. Optimal Transport (OT) offers a different route by comparing the trial and target populations directly in covariate space. We develop RICOT, a Riesz-calibrated OT procedure transporting treatment effects to a treated target population. We consider a semi-unbalanced OT with entropic regularization where the source marginals are relaxed. We show that the uncalibrated OT introduces a bias which does not shrink with increasing sample size. RICOT removes this bias by imposing calibration equations directly within the transport problem. With a growing calibration sieve, the calibrated weight consistently estimates the target-to-trial density ratio, equivalently the Riesz representer of the target expectation functional, even when the entropic and source-relaxation parameters remain fixed and positive. Combined with outcome regression, the resulting estimator is doubly robust and attains the semiparametric efficiency bound under suitable rate conditions. Its variance is estimated directly from the influence function, without resampling or repeated OT optimization. Simulations show low bias and near-nominal coverage across a range of overlap and misspecification settings, including settings in which sampling-score methods perform poorly. We illustrate RICOT in a real-world application involving a rare progressive cardiomyopathy, comparing conventional IPW and AIPW estimators with our OT-based IPW and doubly robust estimators for transporting the randomized treatment effect to a real-world population receiving the same treatment.

stat.ME

Integrating RCTs, RWD, AI/ML and Statistics: Next-Generation Evidence Synthesis

Randomized controlled trials (RCTs)have been the cornerstone of clinical evidence; however, their cost, duration, and restrictive eligibility criteria limit power and external validity. Studies using real-world data (RWD), historically considered less reliable for establishing causality, are now recognized as an important source of real-world evidence (RWE). In parallel, artificial intelligence and machine learning (AI/ML) are increasingly used throughout the drug development process, providing scalability and flexibility but also presenting challenges in interpretability and statistical rigor. This Perspective argues that the future of evidence generation will not depend on RCTs versus RWD, or statistics versus AI/ML, but on their principled integration under a statistical evidence framework that clarifies estimands, evaluates data fitness, controls bias, quantifies uncertainty, and determines when evidence is strong enough to support decisions. Building on the Causal Roadmap for high-quality real-world evidence, we present a six-step statistical roadmap for integrative evidence synthesis and organize the discussion around five core questions that statisticians must confront: when RWD is fit for causal use and when it is not; what AI genuinely contributes across the evidence lifecycle; what remains distinctly statistical and indispensable; how to build trustworthy end-to-end evidence systems in pharmaceutical, regulatory, and industry settings; and how statistical training should evolve.

stat.ME

Risk-inclusive Contextual Bandits for Early Phase Clinical Trials

Early-phase clinical trials face the challenge of selecting optimal drug doses that balance safety and efficacy due to uncertain dose-response relationships and varied participant characteristics. Traditional randomized dose allocation often exposes participants to sub-optimal doses by not considering individual covariates, necessitating larger sample sizes and prolonging drug development. This paper introduces a risk-inclusive contextual bandit algorithm that utilizes multi-arm bandit (MAB) strategies to optimize dosing through participant-specific data integration. By combining two separate Thompson samplers, one for efficacy and one for safety, the algorithm enhances the balance between efficacy and safety in dose allocation. The effect sizes are estimated with a generalized version of asymptotic confidence sequences (AsympCS), offering a uniform coverage guarantee for sequential causal inference over time. The validity of AsympCS is also established in the MAB setup with a possibly mis-specified model. The empirical results demonstrate the strengths of this method in optimizing dose allocation compared to randomized allocations and traditional contextual bandits focused solely on efficacy. Moreover, an application on real data generated from a recent Phase IIb study aligns with actual findings.

stat.ME

Proactive Anomaly Screen for Multiple Endpoints Using Bayesian Latent Class Modeling: A k-Step Ahead Approach

In clinical trials, ensuring the quality and validity of data for downstream analysis and results is paramount, thus necessitating thorough data monitoring. This typically involves employing edit checks and manual queries during data collection. Edit checks consist of straightforward schemes programmed into relational databases, though they lack the capacity to assess data intelligently. In contrast, manual queries are initiated by data managers who manually scrutinize the collected data, identifying discrepancies needing clarification or correction. Manual queries pose significant challenges, particularly when dealing with large-scale data in late-phase clinical trials. Moreover, they are reactive rather than predictive, meaning they address issues after they arise rather than preemptively preventing errors. In this paper, we propose a joint model for multiple endpoints, focusing on primary and key secondary measures, using a Bayesian latent class approach. This model incorporates adjustments for risk monitoring factors, enabling proactive, $k$-step ahead, detection of conflicting or anomalous patterns within the data. Furthermore, we develop individualized dynamic predictions at consecutive time-points to identify potential anomalous values based on observed data. This analysis can be integrated into electronic data capture systems to provide objective alerts to stakeholders. We present simulation results and demonstrate effectiveness of this approach with real-world data.

stat.OT

Extent of Safety Database in Pediatric Drug Development: Types of Assessment, Analytical Precision, and Pathway for Extrapolation through On-Target Effects

Pediatric patients should have access to medicines that have been appropriately evaluated for safety and efficacy. Given this goal of revised labelling, the adequacy of the pediatric clinical development plan and resulting safety database must inform a favorable benefit-risk assessment for the intended use of the medicinal product. While extrapolation from adults can be used to support efficacy of drugs in children, there may be a reluctance to use the same approach in safety assessments, wiping out potential gains in trial efficiency through a reduction of sample size. To address this reluctance, we explore safety review in pediatric trials, including factors affecting these data, specific types of safety assessments, and precision on the estimation of event rates for specific adverse events (AEs) that can be achieved. In addition, we discuss the assessments which can provide a benchmark for the use of extrapolation of safety that focuses on on-target effects. Finally, we explore a unified approach for understanding precision using Bayesian approaches as the most appropriate methodology to describe/ascertain risk in probabilistic terms for the estimate of the event rate of specific AEs.

stat.AP

Composite Likelihoods with Bounded Weights in Extrapolation of Data

Among many efforts to facilitate timely access to safe and effective medicines to children, increased attention has been given to extrapolation. Loosely, it is the leveraging of conclusions or available data from adults or older age groups to draw conclusions for the target pediatric population when it can be assumed that the course of the disease and the expected response to a medicinal product would be sufficiently similar in the pediatric and the reference population. Extrapolation then can be characterized as a statistical mapping of information from the reference (adults or older age groups) to the target pediatric population. The translation, or loosely mapping of information, can be through a composite likelihood approach where the likelihood of the reference population is weighted by exponentiation and that this exponent is related to the value of the mapped information in the target population. The weight is bounded above and below recognizing the fact that similarity (of the disease and the expected response) is still valid despite variability of response between the cohorts. Maximum likelihood approaches are then used for estimation of parameters and asymptotic theory is used to derive distributions of estimates for use in inference. Hence, the estimation of effects in the target population borrows information from reference population. In addition, this manuscript also talks about how this method is related to the Bayesian statistical paradigm.

stat.ME