SearcharxivSearch

arXiv subjects

Satrajit Roychoudhury

Publications and source records attributed to Satrajit Roychoudhury.

17 recordsLinked to original sources

CARB: A Covariate-Assessed Robust Borrowing Strategy with Literature-Informed Prior Weights for External Data

Borrowing external control data can improve the efficiency of clinical trials, particularly when patient accrual is difficult. A persistent challenge is how to prespecify the degree of borrowing systematically and transparently. In practice, prior weights are often selected heuristically or calibrated through simulation to achieve desired operating characteristics, making them difficult to justify scientifically. Moreover, patient-level covariate data from external sources are rarely available when the new trial is designed, precluding patient-level adjustment methods. We propose CARB (Covariate-Assessed Robust Borrowing), a framework that formalizes prior-weight specification as a design-stage assessment of baseline compatibility. Using only aggregate information, CARB quantifies discrepancies in prespecified baseline covariates between the new trial and each external source, without using outcome data from the new trial. A prespecified mapping translates the resulting dissimilarity measure into a source-specific prior weight on the exchangeable component of a robust borrowing model. Simulation studies show that CARB reduces bias and type I error inflation relative to fixed borrowing under observed and unobserved incompatibility, while improving efficiency when external controls are compatible. An application to advanced melanoma trials illustrates covariate-informed borrowing from multiple historical sources. An apparent discrepancy in reported baseline characteristics serves as a warning signal that reduces borrowing. CARB provides a transparent and reproducible way to borrow cautiously when patient-level covariate data from external sources are unavailable.

stat.ME

Dynamic Prediction of the Target Survival Time in Metastatic Solid Tumor Cancer Clinical Trials

Overall survival (OS) is the gold standard for assessing patient benefit and cost-effectiveness of new cancer drugs. However, it is often difficult to use OS as the primary endpoint in randomized clinical trials (RCTs) for patients with metastatic cancer due to multiple reasons. In recent years, progression-free survival (PFS) has increasingly been used as the primary endpoint in metastatic cancer RCTs to accelerate development. However, regulatory authorities often seek mature OS data for approval. Therefore, it is critical to determine the target time when OS data are expected to be mature for reliable statistical inference. Motivated by an advanced renal cell carcinoma (RCC) clinical trial, we develop and investigate different prediction models leveraging information from disease progression to improve target OS prediction times. We propose a multivariate joint modeling approach considering components of progression and OS and extend three models commonly used for association to be used for OS prediction. To the best of our knowledge, this is the first comprehensive statistical study exploring the prediction of OS using different levels of information on disease progression and illustrating these models using a real, complex dataset. Our findings have significant implications for OS prediction.

stat.ME

Learning from Literature: Integrating LLMs and Bayesian Hierarchical Modeling for Oncology Trial Design

Designing modern oncology trials requires synthesizing evidence from prior studies to inform hypothesis generation and sample size determination. Trial designs based on incomplete or imprecise summaries can lead to misspecified hypotheses and underpowered studies, resulting in false positive or negative conclusions. To address this challenge, we developed LEAD-ONC (Literature to Evidence for Analytics and Design in Oncology), an AI-assisted framework that transforms published clinical trial reports into quantitative, design-relevant evidence. Given expert-curated trial publications that meet prespecified eligibility criteria, LEAD-ONC uses large language models to extract baseline characteristics and reconstruct individual patient data from Kaplan-Meier curves, followed by Bayesian hierarchical modeling to generate predictive survival distributions for a prespecified target trial population. We demonstrate the framework using five phase III trials in first-line non-small-cell lung cancer evaluating PD-1 or PD-L1 inhibitors with or without CTLA-4 blockade. Clustering based on baseline characteristics identified three clinically interpretable populations defined by histology. For a prospective randomized trial in the mixed-histology population comparing mono versus dual immune checkpoint inhibition, LEAD-ONC projected a modest median overall survival difference of 2.8 months (95 percent credible interval -2.0 to 7.6) and an estimated probability of at least a 3-month benefit of approximately 0.45. As LEAD-ONC remains under active development, these results are intended as preliminary demonstrations of the frameworks potential to support evidence-driven oncology trial design rather than definitive clinical conclusions.

stat.AP

On regional treatment effect assessment using robust MAP priors

Bayesian dynamic borrowing has become an increasingly important tool for evaluating the consistency of regional treatment effects which is a key requirement for local regulatory approval of a new drug. It helps increase the precision of regional treatment effect estimate when regional and global data are similar, while guarding against potential bias when they differ. In practice, the two-component mixture prior, of which one mixture component utilizes the power prior to incorporate external data, is widely used. It allows convenient prior specification, analytical posterior computation, and fast evaluation of operating characteristics. Though the robust meta-analytical-predictive (MAP) prior is broadly used with multiple external data sources, it remains underutilized for regional treatment effect assessment (typically only one external data source is available) due to its inherit complexity in prior specification and posterior computation. In this article, we illustrate the applicability of the robust MAP prior in the regional treatment effect assessment by developing a closed-form approximation for its posterior distribution while leveraging its relationship with the power prior. The proposed methodology substantially reduces the computational burden of identifying prior parameters for desired operating characteristics. Moreover, we have demonstrated that the MAP prior is an attractive choice to construct the informative component of the mixture prior compared to the power prior. The advantage can be explained through a Bayesian hypothesis testing perspective. Using a real-world example, we illustrate how our proposed method enables efficient and transparent development of a Bayesian dynamic borrowing design to show regional consistency.

stat.ME

Evaluating treatment effects on longitudinal outcomes with attrition due to death: Methods for a two-dimentional estimand with a case study in Quality of Life

When longitudinal outcomes are evaluated in mortal populations, their non-existence after death complicates the analysis and its causal interpretation. Where popular methods often merge longitudinal outcome and survival into one scale or otherwise try to circumvent the problem of mortality, some highly relevant questions require survival to be acknowledged as a unique condition. "\textit{What are my chances of survival}" and "\textit{What can I expect for my condition while still alive}" reflect the intrinsically two-dimensional outcome of survival and longitudinal outcome while-alive. We define a two-dimensional causal while-alive estimand for a point exposure and compare two methods for estimation in an observational setting. Regression-Standardization models survival and the observed longitudinal outcome before standardizing the latter to a target population weighted by its estimated survival. Alternatively, Inverse Probability of Treatment and Censoring Weighting weights the observed outcomes twice, to account for censoring and differences in baseline-case-mix. Both approaches rely on the same causal identification assumptions, but require different models to be correctly specified. With its potential to extrapolate, Regression-Standardization is more efficient when all assumptions are met. We show finite sample performance in a simulation study and apply the methods to a case study on quality of life in oncology.

stat.ME

A robust score test in g-computation for covariate adjustment in randomized clinical trials leveraging different variance estimators via influence functions

G-computation has become a widely used robust method for estimating unconditional (marginal) treatment effects with covariate adjustment in the analysis of randomized clinical trials. Statistical inference in this context typically relies on the Wald test or Wald interval, which can be easily implemented using a consistent variance estimator. However, existing literature suggests that when sample sizes are small or when parameters of interest are near boundary values, Wald-based methods may be less reliable due to type I error rate inflation and insufficient interval coverage. In this article, we propose a robust score test for g-computation estimators in the context of two-sample treatment comparisons. The proposed test is asymptotically valid under simple and stratified (biased-coin) randomization schemes, even when regression models are misspecified. These test statistics can be conveniently computed using existing variance estimators, and the corresponding confidence intervals have closed-form expressions, making them convenient to implement. Through extensive simulations, we demonstrate the superior finite-sample performance of the proposed method. Finally, we apply the proposed method to reanalyze a completed randomized clinical trial. The new analysis using our proposed score test achieves statistical significance, whilst reducing the issue of type I error inflation.

stat.ME

Study Duration Prediction for Clinical Trials with Time-to-Event Endpoints Using Mixture Distributions Accounting for Heterogeneous Population

In the era of precision medicine, more and more clinical trials are now driven or guided by biomarkers, which are patient characteristics objectively measured and evaluated as indicators of normal biological processes, pathogenic processes, or pharmacologic responses to therapeutic interventions. With the overarching objective to optimize and personalize disease management, biomarker-guided clinical trials increase the efficiency by appropriately utilizing prognostic or predictive biomarkers in the design. However, the efficiency gain is often not quantitatively compared to the traditional all-comers design, in which a faster enrollment rate is expected (e.g. due to no restriction to biomarker positive patients) potentially leading to a shorter duration. To accurately predict biomarker-guided trial duration, we propose a general framework using mixture distributions accounting for heterogeneous population. Extensive simulations are performed to evaluate the impact of heterogeneous population and the dynamics of biomarker characteristics and disease on the study duration. Several influential parameters including median survival time, enrollment rate, biomarker prevalence and effect size are identitied. Re-assessments of two publicly available trials are conducted to empirically validate the prediction accuracy and to demonstrate the practical utility. The R package \emph{detest} is developed to implement the proposed method and is publicly available on CRAN.

stat.ME

Duration of and time to response in oncology clinical trials from the perspective of the estimand framework

Duration of response (DOR) and time to response (TTR) are typically evaluated as secondary endpoints in early-stage clinical studies in oncology when efficacy is assessed by the best overall response (BOR) and presented as the overall response rate (ORR). Despite common use of DOR and TTR in particular in single-arm studies, the definition of these endpoints and the questions they are intended to answer remain unclear. Motivated by the estimand framework, we present relevant scientific questions of interest for DOR and TTR and propose corresponding estimand definitions. We elaborate on how to deal with relevant intercurrent events which should follow the same considerations as implemented for the primary response estimand. A case study in mantle cell lymphoma illustrates the implementation of relevant estimands of DOR and TTR. We close the paper with practical recommendations to implement DOR and TTR in clinical study protocols.

stat.ME

The Predictive Individual Effect for Survival Data

The call for patient-focused drug development is loud and clear, as expressed in the 21st Century Cures Act and in recent guidelines and initiatives of regulatory agencies. Among the factors contributing to modernized drug development and improved health-care activities are easily interpretable measures of clinical benefit. In addition, special care is needed for cancer trials with time-to-event endpoints if the treatment effect is not constant over time. We propose the predictive individual effect which is a patient-centric and tangible measure of clinical benefit under a wide variety of scenarios. It can be obtained by standard predictive calculations under a rank preservation assumption that has been used previously in trials with treatment switching. We discuss four recent Oncology trials that cover situations with proportional as well as non-proportional hazards (delayed treatment effect or crossing of survival curves). It is shown that the predictive individual effect offers valuable insights beyond p-values, estimates of hazard ratios or differences in median survival. Compared to standard statistical measures, the predictive individual effect is a direct, easily interpretable measure of clinical benefit. It facilitates communication among clinicians, patients, and other parties and should therefore be considered in addition to standard statistical results.

stat.AP

Principal Stratum Strategy: Potential Role in Drug Development

A randomized trial allows estimation of the causal effect of an intervention compared to a control in the overall population and in subpopulations defined by baseline characteristics. Often, however, clinical questions also arise regarding the treatment effect in subpopulations of patients, which would experience clinical or disease related events post-randomization. Events that occur after treatment initiation and potentially affect the interpretation or the existence of the measurements are called {\it intercurrent events} in the ICH E9(R1) guideline. If the intercurrent event is a consequence of treatment, randomization alone is no longer sufficient to meaningfully estimate the treatment effect. Analyses comparing the subgroups of patients without the intercurrent events for intervention and control will not estimate a causal effect. This is well known, but post-hoc analyses of this kind are commonly performed in drug development. An alternative approach is the principal stratum strategy, which classifies subjects according to their potential occurrence of an intercurrent event on both study arms. We illustrate with examples that questions formulated through principal strata occur naturally in drug development and argue that approaching these questions with the ICH E9(R1) estimand framework has the potential to lead to more transparent assumptions as well as more adequate analyses and conclusions. In addition, we provide an overview of assumptions required for estimation of effects in principal strata. Most of these assumptions are unverifiable and should hence be based on solid scientific understanding. Sensitivity analyses are needed to assess robustness of conclusions.

stat.AP

Robust Design and Analysis of Clinical Trials With Non-proportional Hazards: A Straw Man Guidance from a Cross-pharma Working Group

Loss of power and clear description of treatment differences are key issues in designing and analyzing a clinical trial where non-proportional hazard is a possibility. A log-rank test may be very inefficient and interpretation of the hazard ratio estimated using Cox regression is potentially problematic. In this case, the current ICH E9 (R1) addendum would suggest designing a trial with a clinically relevant estimand, e.g., expected life gain. This approach considers appropriate analysis methods for supporting the chosen estimand. However, such an approach is case specific and may suffer lack of power for important choices of the underlying alternate hypothesis distribution. On the other hand, there may be a desire to have robust power under different deviations from proportional hazards. Also, we would contend that no single number adequately describes treatment effect under non-proportional hazards scenarios. The cross-pharma working group has proposed a combination test to provide robust power under a variety of alternative hypotheses. These can be specified for primary analysis at the design stage and methods appropriately accounting for combination test correlations are efficient for a variety of scenarios. We have provided design and analysis considerations based on a combination test under different non-proportional hazard types and present a straw man proposal for practitioners. The proposals are illustrated with real life example and simulation.

stat.AP

Estimands in Hematologic Oncology Trials

The estimand framework included in the addendum to the ICH E9 guideline facilitates discussions to ensure alignment between the key question of interest, the analysis, and interpretation. Therapeutic knowledge and drug mechanism play a crucial role in determining the strategy and defining the estimand for clinical trial designs. Clinical trials in patients with hematological malignancies often present unique challenges for trial design due to complexity of treatment options and existence of potential curative but highly risky procedures, e.g. stem cell transplant or treatment sequence across different phases (induction, consolidation, maintenance). Here, we illustrate how to apply the estimand framework in hematological clinical trials and how the estimand framework can address potential difficulties in trial result interpretation. This paper is a result of a cross-industry collaboration to connect the International Conference on Harmonisation (ICH) E9 addendum concepts to applications. Three randomized phase 3 trials will be used to consider common challenges including intercurrent events in hematologic oncology trials to illustrate different scientific questions and the consequences of the estimand choice for trial design, data collection, analysis, and interpretation. Template language for describing estimand in both study protocols and statistical analysis plans is suggested for statisticians' reference.

q-bio.OT

Bayesian leveraging of historical control data for a clinical trial with time-to-event endpoint

The recent 21st Century Cures Act propagates innovations to accelerate the discovery, development, and delivery of 21st century cures. It includes the broader application of Bayesian statistics and the use of evidence from clinical expertise. An example of the latter is the use of trial-external (or historical) data, which promises more efficient or ethical trial designs. We propose a Bayesian meta-analytic approach to leveraging historical data for time-to-event endpoints, which are common in oncology and cardiovascular diseases. The approach is based on a robust hierarchical model for piecewise exponential data. It allows for various degrees of between trial-heterogeneity and for leveraging individual as well as aggregate data. An ovarian carcinoma trial and a non-small-cell cancer trial illustrate methodological and practical aspects of leveraging historical data for the analysis and design of time-to-event trials.

stat.AP

Alternative Analysis Methods for Time to Event Endpoints under Non-proportional Hazards: A Comparative Analysis

The log-rank test is most powerful under proportional hazards (PH). In practice, non-PH patterns are often observed in clinical trials, such as in immuno-oncology; therefore, alternative methods are needed to restore the efficiency of statistical testing. Three categories of testing methods were evaluated, including weighted log-rank tests, Kaplan-Meier curve-based tests (including weighted Kaplan-Meier and Restricted Mean Survival Time, RMST), and combination tests (including Breslow test, Lee's combo test, and MaxCombo test). Nine scenarios representing the PH and various non-PH patterns were simulated. The power, type I error, and effect estimates of each method were compared. In general, all tests control type I error well. There is not a single most powerful test across all scenarios. In the absence of prior knowledge regarding the PH or non-PH patterns, the MaxCombo test is relatively robust across patterns. Since the treatment effect changes overtime under non-PH, the overall profile of the treatment effect may not be represented comprehensively based on a single measure. Thus, multiple measures of the treatment effect should be pre-specified as sensitivity analyses to evaluate the totality of the data.

stat.AP

Beyond p-values: a phase II dual-criterion design with statistical significance and clinical relevance

Background: Well-designed phase II trials must have acceptable error rates relative to a pre-specified success criterion, usually a statistically significant p-value. Such standard designs may not always suffice from a clinical perspective because clinical relevance may call for more. For example, proof-of-concept in phase II often requires not only statistical significance but also a sufficiently large effect estimate. Purpose: We propose dual-criterion designs to complement statistical significance with clinical relevance, discuss their methodology, and illustrate their implementation in phase II. Methods: Clinical relevance requires the effect estimate to pass a clinically motivated threshold (the decision value). In contrast to standard designs, the required effect estimate is an explicit design input whereas study power is implicit. The sample size for a dual-criterion design needs careful considerations of the study's operating characteristics (type-I error, power). Results: Dual-criterion designs are discussed for a randomized controlled and a single-arm phase II trial, including decision criteria, sample size calculations, decisions under various data scenarios, and operating characteristics. The designs facilitate GO/NO-GO decisions due to their complementary statistical-clinical criterion. Conclusion: To improve evidence-based decision-making, a formal yet transparent quantitative framework is important. Dual-criterion designs offer an appealing statistical-clinical compromise, which may be preferable to standard designs if evidence against the null hypothesis alone does not suffice for an efficacy claim.

stat.AP

On the structure of a family of probability generating functions induced by shock models

We explore conditions for a class of functions defined via an integral representation to be a probability generating function of some positive integer valued random variable. Interest in and research on this question is motivated by an apparently surprising connection between a family of classic shock models due to Esary et. al. (1973) and the negatively aging nonparametric notion of ``strongly decreasing failure rate'' (SDFR) introduced by Bhattacharjee (2005). A counterexample shows that there exist probability generating functions with our integral representation which are not discrete SDFR, but when used as shock resistance probabilities can give rise to a SDFR survival distribution in continuous time.

math.ST