SearcharxivSearch

arXiv subjects

Thomas Jaki

Publications and source records attributed to Thomas Jaki.

At least 19 recordsLinked to original sources

Introducing precision-weighted bias as a performance measure to inform the inclusion of adaptive designs in meta-analysis

We propose a novel, intuitive measure of statistical performance: precision-weighted bias. Precision-weighted bias is defined as the unconditional bias of an estimator weighted by the degree of information (precision) it contains. Current guidelines, such as GRADE and CONSORT, often view the potential for increased bias in adaptive designs as a deterrent for the inclusion of such designs in systematic reviews. However, we demonstrate that the bias in a common-effect meta-analysis is approximately equal to the precision-weighted average of the precision-weighted biases of its constituent studies, rather than of their unweighted unconditional biases. Through simulation studies, we show that while adaptive designs may exhibit unweighted bias, they frequently have zero precision-weighted bias. Consequently, including these designs often results in a negligible change to the overall meta-analysis bias. These results suggest that precision-weighted bias is a superior indicator for determining whether to include an adaptive design in a meta-analysis. We recommend that precision-weighted bias be used as a standard complement to unweighted unconditional and conditional bias in simulation studies to support more inclusive and accurate evidence synthesis.

stat.ME

Operationalizing Allocation Probability Tests: Practical Guidance on Optimized Implementation for Power and Robustness

Recently, a new testing approach for response-adaptive clinical trials was proposed based on the allocation probabilities (AP) rather than the outcome data. While original work on the AP test focused on binary and normal endpoints and demonstrated that significant efficiency gains are possible, many critical questions remain open regarding its practical implementation and upper limits. In this work, rather than simply proposing novel statistics, we seek to understand the maximum gain that can be obtained with the AP test by optimizing how these probabilities are used to define the test statistic. We expand the method's practical utility by applying it to survival endpoints (exponential distributions) and introducing a rigorous strategy for selecting the null hypothesis to properly calibrate type I error. Our simulation studies reveal that by optimizing the functional form of the AP test, investigators can achieve a substantial increase in power, approaching the theoretical maximum, without sacrificing the patient outcome goals of the design. Furthermore, we explicitly compare the method to a standard Bayesian decision rule, finding that the optimized AP test significantly outperforms traditional frequentist tests while maintaining strict error control. This work provides a missing practical framework for implementing robust and optimized AP tests in complex response-adaptive settings.

stat.ME

Robust Bayesian Sequential Borrowing for Multi-Population Clinical Programmes

We introduce Robust Bayesian Sequential Borrowing (RBSB), a framework for extrapolating evidence across adjacent subgroups in multi-population clinical programmes where studies are conducted in sequence and populations are ordered by clinical proximity. Conventional approaches weight all historical sources uniformly or exclude distant populations entirely, failing to reflect the natural gradient of similarity in such programmes. RBSB encodes the programme order through path-dependent borrowing via robust mixture priors that combine an informative component with a unit-information component to guard against prior-data conflict. Posterior weights, derived in closed form from marginal likelihood ratios, provide transparent dynamic attenuation when heterogeneity arises between sequential populations. The framework supports prospective evaluation of Bayesian Type I error, power, and extends naturally to assurance at both the study and programme level. Simulation studies demonstrate superior false-positive control relative to full pooling, while preserving substantial efficiency gains over standalone analyses. A case study of the START trial illustrates the approach across adult, adolescent, and paediatric populations. RBSB offers a practical, regulator-aligned method for disciplined evidence borrowing that exploits temporal and biological proximity while preventing implausible extrapolation across distant populations.

stat.ME

Evaluating Predictive Modeling Strategies for Predicting Individual Treatment Effects in Precision Medicine

Precision medicine seeks to match patients with treatments that produce the greatest benefit. The Predicted Individual Treatment Effect (PITE)-the difference between predicted outcomes under treatment and control-quantifies this benefit but is difficult to estimate due to unobserved counterfactuals, high dimensionality, and complex interactions. We compared 30+ modeling strategies, including penalized and projection-based methods, flexible learners, and tree-ensembles, using a structured simulation framework varying sample size, dimensionality, multicollinearity, and interaction complexity. Performance was measured using root mean squared error (RMSE) for prediction accuracy and directional accuracy (DIR) for correctly classifying benefit versus harm. Internal validation produced optimistic estimates, whereas external validation with distributional shifts and higher-order interactions more clearly revealed model weaknesses. Penalized and projection-based approaches-ridge, lasso, elastic net, partial least squares (PLS), and principal components regression (PCR)-consistently achieved strong RMSE and DIR performance. Flexible learners excelled only under strong signals and sufficient sample sizes. Results highlight robust linear/projection defaults and the necessity of rigorous external validation.

stat.AP

Optimal weighted tests for replication studies and the two-trials rule

Replication studies for scientific research are an important part of ensuring the reliability and integrity of experimental findings. In the context of clinical trials, the concept of replication has been formalised by the 'two-trials' rule, where two pivotal studies are required to show positive results before a drug can be approved. In experiments testing multiple hypotheses simultaneously, control of the overall familywise error rate (FWER) is additionally required in many contexts. The well-known Bonferroni procedure controls the FWER, and a natural extension is to introduce weights into this procedure to reflect the a-priori importance of hypotheses or to maximise some measure of the overall power of the experiment. In this paper, we consider analysing a replication study using an optimal weighted Bonferroni procedure, with the weights based on the results of the original study that is being replicated and the optimality criterion being to maximise the disjunctive power of the trial (the power to reject at least one non-null hypothesis). We show that using the proposed procedure can lead to a substantial increase in the disjunctive power of the replication study, and is robust to changes in the effect sizes between the two studies.

stat.ME

A comparison of approaches to incorporate patient-selected and patient-ranked outcomes in clinical trials

A key aspect of patient-focused drug development is identifying and measuring outcomes that are important to patients in clinical trials. Many medical conditions affect multiple symptom domains, and a consensus approach to determine the relative importance of the associated multiple outcomes ignores the heterogeneity in individual patient preferences. Patient-selected outcomes offer one way to incorporate individual patient preferences, as proposed in recent regulatory guidance for the treatment for migraine, where each patient selects their most bothersome migraine-associated symptom in addition to pain. Patient-ranked outcomes have also recently been proposed, which go further and consider the full ranking of the relative importance of all the outcomes. This can be assessed using a composite DOOR (Desirability of Outcome Ranking) endpoint. In this paper, we compare the advantages and disadvantages of using patient-selected versus patient-ranked outcomes in the context of a two-arm randomised controlled trial for multiple sclerosis. We compare the power and type I error rate by simulation, and discuss several other important considerations when using the two approaches.

stat.ME

Using joint models in phase I dose-finding designs in oncology: considerations for frequentist approaches

Dose-finding trials for oncology studies are traditionally designed to assess safety in the early stages of drug development. With the rise of molecularly targeted therapies and immuno-oncology compounds, biomarker-driven approaches have gained significant importance. In this paper, we propose a novel approach that incorporates multiple values of a predictive biomarker to assist in evaluating binary toxicity outcomes using the factorization of a joint model in phase I dose-finding oncology trials. The proposed joint model framework, which utilizes additional repeated biomarker values as an early predictive marker for potential toxicity, is compared to the likelihood-based continual reassessment method (CRM) using only binary toxicity data, across various dose-toxicity relationship scenarios. Our findings highlight a critical limitation of likelihood-based approaches in early-phase dose-finding studies with small sample sizes: estimation challenges that have been previously overlooked in the phase I dose-escalation setting. We explore potential remedies to address these challenges and emphasize the appropriate use of likelihood-based methods. Simulation results demonstrate that the proposed joint model framework, by integrating biomarker information, can alleviate estimation problems in the the likelihood-based continual reassessment method (CRM) and improve the proportion of correct selection. However, we highlight that the inherent data limitations in early-phase dose-finding studies remain a significant challenge that cannot fully be overcomed in the frequentist framework.

stat.ME

A Bayesian Additive Regression Trees Model for zero and one inflated data for Predicting Individual Treatment Effects in Alcohol Use Disorder Trials

Alcohol Use Disorder (AUD) treatment presents high individual-level heterogeneity, with outcomes ranging from complete abstinence to persistent heavy drinking. This variability-driven by complex behavioral, social, and environmental factors-poses major challenges for treatment evaluation and individualized decision-making. In particular, accurately modeling bounded semicontinuous outcomes and estimating predictive individual treatment effects (PITEs) remains methodologically demanding. For the pre-registered PITE analysis of Project MATCH, we developed HOBZ-BART, a novel Bayesian nonparametric model tailored for semicontinuous outcomes concentrated at clinically meaningful boundary values (0 and 1). The model decomposes the outcome into three components-abstinence, partial drinking, and persistent use-via a sequential hurdle structure, offering interpretability aligned with clinical reasoning. A shared Bayesian Additive Regression Tree (BART) ensemble captures nonlinear effects and covariate interactions across components, while a scalable Beta-likelihood approximation enables efficient, conjugate-friendly posterior computation. Through extensive simulations we demonstrate that HOBZ-BART outperforms traditional zero-one inflated Beta (ZOIB) model in predictive accuracy, computational efficiency, and PITE estimation. We then present the primary PITE analysis of the MATCH trial using HOBZ-BART which enables clinically meaningful comparisons of Cognitive Behavioral Therapy (CBT), Motivational Enhancement Therapy (MET), and Twelve Step Facilitation (TSF), offering personalized treatment insights. HOBZ-BART combines statistical rigor with clinical interpretability, addressing a critical need in addiction research for models that support individualized, data-driven care.

stat.AP

A multi-arm multi-stage design for trials with all pairwise testing

Multi-arm multi-stage (MAMS) trials have gained popularity to enhance the efficiency of clinical trials, potentially reducing both duration and costs. This paper focuses on designing MAMS trials where no control treatment exists. This can arise when multiple standard treatments are already established or no treatment is available for a severe disease, making it unethical to withhold a potentially helpful option. The proposed design incorporates interim analyses to allow early termination of notably worst treatments and stops the trial entirely if all remaining treatments are performing similarly. The proposed design controls the familywise error rate (FWER) for all pairwise comparisons and provides the conditions guaranteeing FWER control in the strong sense. The FWER and power are used to calculate both the stopping boundaries and the sample size required. Analytic solutions to compute the expected sample size are also derived. A trial motivated by a study conducted in sepsis, where there was no control treatment, is shown. The multi-arm multi-stage all pairwise (MAMSAP) design proposed here is compared to multiple different approaches. For the trial studied, the proposed method yields the lowest required maximum and expected sample size when controlling the FWER and power at the desired levels.

stat.ME

Joint TITE-CRM for Dual Agent Dose Finding Studies

Dual agent dose-finding trials study the effect of a combination of more than one agent, where the objective is to find the Maximum Tolerated Dose Combination (MTC), the combination of doses of the two agents that is associated with a pre-specified risk of being unsafe. In a Phase I/II setting, the objective is to find a dose combination that is both safe and active, the Optimal Biological Dose (OBD), that optimizes a criterion based on both safety and activity. Since Oncology treatments are typically given over multiple cycles, both the safety and activity outcome can be considered as late-onset, potentially occurring in the later cycles of treatment. This work proposes two model-based designs for dual-agent dose finding studies with late-onset activity and late-onset toxicity outcomes, the Joint TITE-POCRM and the Joint TITE-BLRM. Their performance is compared alongside a model-assisted comparator in a comprehensive simulation study motivated by a real trial example, with an extension to consider alternative sized dosing grids. It is found that both model-based methods outperform the model-assisted design. Whilst on average the two model-based designs are comparable, this comparability is not consistent across scenarios.

stat.AP

Confidence intervals for adaptive trial designs I: A methodological review

Regulatory guidance notes the need for caution in the interpretation of confidence intervals (CIs) constructed during and after an adaptive clinical trial. Conventional CIs of the treatment effects are prone to undercoverage (as well as other undesirable properties) in many adaptive designs, because they do not take into account the potential and realised trial adaptations. This paper is the first in a two-part series that explores CIs for adaptive trials. It provides a comprehensive review of the methods to construct CIs for adaptive designs, while the second paper illustrates how to implement these in practice and proposes a set of guidelines for trial statisticians. We describe several classes of techniques for constructing CIs for adaptive clinical trials, before providing a systematic literature review of available methods, classified by the type of adaptive design. As part of this, we assess, through a proposed traffic light system, which of several desirable features of CIs (such as achieving nominal coverage and consistency with the hypothesis test decision) each of these methods holds.

stat.ME

Confidence intervals for adaptive trial designs II: Case study and practical guidance

In adaptive clinical trials, the conventional confidence interval (CI) for a treatment effect is prone to undesirable properties such as undercoverage and potential inconsistency with the final hypothesis testing decision. Accordingly, as is stated in recent regulatory guidance on adaptive designs, there is the need for caution in the interpretation of CIs constructed during and after an adaptive clinical trial. However, it may be unclear which of the available CIs in the literature are preferable. This paper is the second in a two-part series that explores CIs for adaptive trials. Part I provided a methodological review of approaches to construct CIs for adaptive designs. In this paper (part II), we present an extended case study based around a two-stage group sequential trial, including a comprehensive simulation study of the proposed CIs for this setting. This facilitates an expanded description of considerations around what makes for an effective CI procedure following an adaptive trial. We show that the CIs can have notably different properties. Finally, we propose a set of guidelines for researchers around the choice of CIs and the reporting of CIs following an adaptive design.

stat.ME

Making all pairwise comparisons in multi-arm clinical trials without control treatment

The standard paradigm for confirmatory clinical trials is to compare experimental treatments with a control, for example the standard of care or a placebo. However, it is not always the case that a suitable control exists. Efficient statistical methodology is well studied in the setting of randomised controlled trials. This is not the case if one wishes to compare several experimental with no control arm. We propose hypothesis testing methods suitable for use in such a setting. These methods are efficient, ensuring the error rate is controlled at exactly the desired rate with no conservatism. This in turn yields an improvement in power when compared with standard methods one might otherwise consider using, such as a Bonferroni adjustment. The proposed testing procedure is also highly flexible. We show how it may be extended for use in multi-stage adaptive trials, covering the majority of scenarios in which one might consider the use of such procedures in the clinical trials setting. With such a highly flexible nature, these methods may also be applied more broadly outside of a clinical trials setting.

stat.ME

A fast, flexible simulation framework for Bayesian adaptive designs -- the R package BATSS

The use of Bayesian adaptive designs for randomised controlled trials has been hindered by the lack of software readily available to statisticians. We have developed a new software package (Bayesian Adaptive Trials Simulator Software - BATSS for the statistical software R, which provides a flexible structure for the fast simulation of Bayesian adaptive designs for clinical trials. We illustrate how the BATSS package can be used to define and evaluate the operating characteristics of Bayesian adaptive designs for various different types of primary outcomes (e.g., those that follow a normal, binary, Poisson or negative binomial distribution) and can incorporate the most common types of adaptations: stopping treatments (or the entire trial) for efficacy or futility, and Bayesian response adaptive randomisation - based on user-defined adaptation rules. Other important features of this highly modular package include: the use of (Integrated Nested) Laplace approximations to compute posterior distributions, parallel processing on a computer or a cluster, customisability, adjustment for covariates and a wide range of available conditional distributions for the response.

stat.CO

How to Add Baskets to an Ongoing Basket Trial with Information Borrowing

Basket trials test a single therapeutic treatment on several patient populations under one master protocol. A desirable adaptive design feature in these studies may be the incorporation of new baskets to an ongoing study. Limited basket sample sizes can cause issues in power and precision of treatment effect estimates which could be amplified in added baskets due to the shortened recruitment time. While various Bayesian information borrowing techniques have been introduced to tackle the issue of small sample sizes, the impact of including new baskets in the trial and into the borrowing model has yet to be investigated. We explore approaches for adding baskets to an ongoing trial under information borrowing and highlight when it is beneficial to add a basket compared to running a separate investigation for new baskets. We also propose a novel calibration approach for the decision criteria that is more robust to false decision making. Simulation studies are conducted to assess the performance of approaches which is monitored primarily through type I error control and precision of estimates. Results display a substantial improvement in power for a new basket when information borrowing is utilized, however, this comes with potential inflation of error rates which can be shown to be reduced under the proposed calibration procedure.

stat.ME

Next generation clinical trials: Seamless designs and master protocols

Background: Drug development is often inefficient, costly and lengthy, yet it is essential for evaluating the safety and efficacy of new interventions. Compared with other disease areas, this is particularly true for Phase II / III cancer clinical trials where high attrition rates and reduced regulatory approvals are being seen. In response to these challenges, seamless clinical trials and master protocols have emerged to streamline the drug development process. Methods: Seamless clinical trials, characterized by their ability to transition seamlessly from one phase to another, can lead to accelerating the development of promising therapies while Master protocols provide a framework for investigating multiple treatment options and patient subgroups within a single trial. Results: We discuss the advantages of these methods through real trial examples and the principals that lead to their success while also acknowledging the associated regulatory considerations and challenges. Conclusion: Seamless designs and Master protocols have the potential to improve confirmatory clinical trials. In the disease area of cancer, this ultimately means that patients can receive life-saving treatments sooner.

stat.ME

Multivariate group sequential tests for global summary statistics

We describe group sequential tests which efficiently incorporate information from multiple endpoints allowing for early stopping at pre-planned interim analyses. We formulate a testing procedure where several outcomes are examined, and interim decisions are based on a global summary statistic. An error spending approach to this problem is defined which allows for unpredictable group sizes and nuisance parameters such as the correlation between endpoints. We present and compare three methods for implementation of the testing procedure including numerical integration, the Delta approximation and Monte Carlo simulation. In our evaluation, numerical integration techniques performed best for implementation with error rate calculations accurate to five decimal places. Our proposed testing method is flexible and accommodates summary statistics derived from general, non-linear functions of endpoints informed by the statistical model. Type 1 error rates are controlled, and sample size calculations can easily be performed to satisfy power requirements.

stat.ME

Bayesian Model Averaging for Partial Ordering Continual Reassessment Methods

Phase I clinical trials are essential to bringing novel therapies from chemical development to widespread use. Traditional approaches to dose-finding in Phase I trials, such as the '3+3' method and the Continual Reassessment Method (CRM), provide a principled approach for escalating across dose levels. However, these methods lack the ability to incorporate uncertainty regarding the dose-toxicity ordering as found in combination drug trials. Under this setting, dose-levels vary across multiple drugs simultaneously, leading to multiple possible dose-toxicity orderings. The Partial Ordering CRM (POCRM) extends to these settings by allowing for multiple dose-toxicity orderings. In this work, it is shown that the POCRM is vulnerable to 'estimation incoherency' whereby toxicity estimates shift in an illogical way, threatening patient safety and undermining clinician trust in dose-finding models. To this end, the Bayesian model averaged POCRM (BMA-POCRM) is proposed. BMA-POCRM uses Bayesian model averaging to take into account all possible orderings simultaneously, reducing the frequency of estimation incoherencies. The effectiveness of BMA-POCRM in drug combination settings is demonstrated through a specific instance of estimate incoherency of POCRM and simulation studies. The results highlight the improved safety, accuracy and reduced occurrence of estimate incoherency in trials applying the BMA-POCRM relative to the POCRM model.

stat.ME