SearcharxivSearch

arXiv subjects

Sarah Zohar

Publications and source records attributed to Sarah Zohar.

15 recordsLinked to original sources

Clustering-Based Outcome Models for Clinical Studies: A Scoping Review

This review provides a systematic overview of methods that combine covariate-based clustering of observational units (patients) with outcome models for clinical studies. We distinguish between informed-cluster models, where the outcome contributes to cluster formation, and agnostic-cluster models, where clustering is performed solely on covariates in a separate first step. Informed-cluster models include product partition models with covariates (PPMx), finite mixtures of regression models (FMR), and cluster-aware supervised learning (CluSL). Agnostic-cluster models encompass two-step procedures using either model-based or algorithmic clustering followed by cluster-specific regression models. Following a systematic search of Web of Science and PubMed, 55 records were identified that propose or evaluate such models. We describe the key models, summarise study characteristics, and present applications from biomedical and public health research. Clustering-based outcome models are particularly relevant for settings with high-dimensional covariates (e.g., biomarker panels and "omics") and heterogeneous patient populations. These models can support risk stratification and we discuss extensions to estimate subgroup-specific treatment effects. They are most valuable when the population is clustered in distinct regions of the covariate space that correspond to different outcome distributions. We discuss applications to rare disease research, covariate adjustment and borrowing from historical data, and subgroup-specific treatment effect estimation in clinical trials.

stat.ME

Latent Neural-ODE for Model-Informed Precision Dosing: Overcoming Structural Assumptions in Pharmacokinetics

Accurate estimation of tacrolimus exposure, quantified by the area under the concentration-time curve (AUC), is essential for precision dosing after renal transplantation. Current practice relies on population pharmacokinetic (PopPK) models based on nonlinear mixed-effects (NLME) methods. However, these models depend on rigid, pre-specified assumptions and may struggle to capture complex, patient-specific dynamics, leading to model misspecification. In this study, we introduce a novel data-driven alternative based on Latent Ordinary Differential Equations (Latent ODEs) for tacrolimus AUC prediction. This deep learning approach learns individualized pharmacokinetic dynamics directly from sparse clinical data, enabling greater flexibility in modeling complex biological behavior. The model was evaluated through extensive simulations across multiple scenarios and benchmarked against two standard approaches: NLME-based estimation and the iterative two-stage Bayesian (it2B) method. We further performed a rigorous clinical validation using a development dataset (n = 178) and a completely independent external dataset (n = 75). In simulation, the Latent ODE model demonstrated superior robustness, maintaining high accuracy even when underlying biological mechanisms deviated from standard assumptions. Regarding experiments on clinical datasets, in internal validation, it achieved significantly higher precision with a mean RMSPE of 7.99% compared with 9.24% for it2B (p < 0.001). On the external cohort, it achieved an RMSPE of 10.82%, comparable to the two standard estimators (11.48% and 11.54%). These results establish the Latent ODE as a powerful and reliable tool for AUC prediction. Its flexible architecture provides a promising foundation for next-generation, multi-modal models in personalized medicine.

stat.ML

In silico clinical trials in drug development: a systematic review

In the context of clinical research, computational models have received increasing attention over the past decades. In this systematic review, we aimed to provide an overview of the role of so-called in silico clinical trials (ISCTs) in medical applications. Exemplary for the broad field of clinical medicine, we focused on in silico (IS) methods applied in drug development, sometimes also referred to as model informed drug development (MIDD). We searched PubMed and ClinicalTrials.gov for published articles and registered clinical trials related to ISCTs. We identified 202 articles and 48 trials, and of these, 76 articles and 19 trials were directly linked to drug development. We extracted information from all 202 articles and 48 clinical trials and conducted a more detailed review of the methods used in the 76 articles that are connected to drug development. Regarding application, most articles and trials focused on cancer and imaging-related research while rare and pediatric diseases were only addressed in 14 articles and 5 trials, respectively. While some models were informed combining mechanistic knowledge with clinical or preclinical (in-vivo or in-vitro) data, the majority of models were fully data-driven, illustrating that clinical data is a crucial part in the process of generating synthetic data in ISCTs. Regarding reproducibility, a more detailed analysis revealed that only 24% (18 out of 76) of the articles provided an open-source implementation of the applied models, and in only 20% of the articles the generated synthetic data were publicly available. Despite the widely raised interest, we also found that it is still uncommon for ISCTs to be part of a registered clinical trial and their application is restricted to specific diseases leaving potential benefits of ISCTs not fully exploited.

q-bio.QM

Straightforward Phase I Dose-Finding Design for Healthy Volunteers Accounting for Surrogate Activity Biomarkers

Conventionally, a first-in-human phase I trial in healthy volunteers aims to confirm the safety of a drug in humans. In such situations, volunteers should not suffer from any safety issues and simple algorithm-based dose-escalation schemes are often used. However, to avoid too many clinical trials in the future, it might be appealing to design these trials to accumulate information on the link between dose and efficacy/activity under strict safety constraints. Furthermore, an increasing number of molecules for which the increasing dose-activity curve reaches a plateau are emerging.In a phase I dose-finding trial context, our objective is to determine, under safety constraints, among a set of doses, the lowest dose whose probability of activity is closest to a given target. For this purpose, we propose a two-stage dose-finding design. The first stage is a typical algorithm dose escalation phase that can both check the safety of the doses and accumulate activity information. The second stage is a model-based dose-finding phase that involves selecting the best dose-activity model according to the plateau location.Our simulation study shows that our proposed method performs better than the common Bayesian logistic regression model in selecting the optimal dose.

stat.AP

An efficient joint model for high dimensional longitudinal and survival data via generic association features

This paper introduces a prognostic method called FLASH that addresses the problem of joint modelling of longitudinal data and censored durations when a large number of both longitudinal and time-independent features are available. In the literature, standard joint models are either of the shared random effect or joint latent class type. Combining ideas from both worlds and using appropriate regularisation techniques, we define a new model with the ability to automatically identify significant prognostic longitudinal features in a high-dimensional context, which is of increasing importance in many areas such as personalised medicine or churn prediction. We develop an estimation methodology based on the EM algorithm and provide an efficient implementation. The statistical performance of the method is demonstrated both in extensive Monte Carlo simulation studies and on publicly available real-world datasets. Our method significantly outperforms the state-of-the-art joint models in predicting the latent class membership probability in terms of the C-index in a so-called ``real-time'' prediction setting, with a computational speed that is orders of magnitude faster than competing methods. In addition, our model automatically identifies significant features that are relevant from a practical perspective, making it interpretable.

stat.ME

Bayesian Framework for Multi-Source Data Integration -- Application to Human Extrapolation From Preclinical Studies

In preclinical investigations, e.g. in in vitro, in vivo and in silico studies, the pharmacokinetic, pharmacodynamic and toxicological characteristics of a drug are evaluated before advancing to first-in-man trial. Usually, each study is analyzed independently and the human dose range does not leverage the knowledge gained from all studies. Taking into account the preclinical data through inferential procedures can be particularly interesting to obtain a more precise and reliable starting dose and dose range. We propose a Bayesian framework for multi-source data integration from preclinical studies results extrapolated to human, which allow to predict the quantities of interest (e.g. the minimum effective dose, the maximum tolerated dose, etc.) in humans. We build an approach, divided in four main steps, based on a sequential parameter estimation for each study, extrapolation to human, commensurability checking between posterior distributions and final information merging to increase the precision of estimation. The new framework is evaluated via an extensive simulation study, based on a real-life example in oncology inspired from the preclinical development of galunisertib. Our approach allows to better use all the information compared to a standard framework, reducing uncertainty in the predictions and potentially leading to a more efficient dose selection.

stat.ME

Coping with Information Loss and the Use of Auxiliary Sources of Data: A Report from the NISS Ingram Olkin Forum Series on Unplanned Clinical Trial Disruptions

Clinical trials disruption has always represented a non negligible part of the ending of interventional studies. While the SARS-CoV-2 (COVID-19) pandemic has led to an impressive and unprecedented initiation of clinical research, it has also led to considerable disruption of clinical trials in other disease areas, with around 80% of non-COVID-19 trials stopped or interrupted during the pandemic. In many cases the disrupted trials will not have the planned statistical power necessary to yield interpretable results. This paper describes methods to compensate for the information loss arising from trial disruptions by incorporating additional information available from auxiliary data sources. The methods described include the use of auxiliary data on baseline and early outcome data available from the trial itself and frequentist and Bayesian approaches for the incorporation of information from external data sources. The methods are illustrated by application to the analysis of artificial data based on the Primary care pediatrics Learning Activity Nutrition (PLAN) study, a clinical trial assessing a diet and exercise intervention for overweight children, that was affected by the COVID-19 pandemic. We show how all of the methods proposed lead to an increase in precision relative to use of complete case data only.

stat.AP

How to improve the quality of comparisons using external control cohorts in single-arm clinical trials?

PURPOSE Providing rapid answers and early acces to patients to innovative treatments without randomized clinical trial (RCT) is growing, with benefit estimated from single-arm trials. This has become common in oncology, impacting the approval pathway of health technology assessment agencies. We aimed to provide some guidance for indirect comparison to external controls to improve the level of evidence following such uncontrolled designs. METHODS We used the illustrative example of blinatumomab, a bispecific antibody for the treatment of B-cell ALL in complete remission (CR) with persistent minimal residual disease (MRD). Its approval relied on a single-arm trial conducted in 86 adults with B-cell ALL in CR, with undetectable MRD after one cycle as the main endpoint. To maximize the validity of indirect comparisons, a 3-step process for incorporating external control data to such single-arm trial data is proposed and detailed, with emphasis on the example. RESULTS The first step includes the definition of estimand, i.e. the treatment effect reflecting the clinical question. The second step relies on the adequate selection of external controls, from previous RCT or real-world data (RWD) obtained from patient cohort, registries, or electronic patient files. The third step consists in chosing the statistical approach targeting the treatment effect of interest, either in the whole population or restricted to the single-arm trial or the external controls, and depending on the available individual-level or aggregrated external data. CONCLUSION Validity of treatment effect derived from indirect comparisons heavily depends on carefull methodological considerations that are included in the proposed 3-step procedure. Because the level of evidence of a well conducted RCT cannot be guaranteed, post-market authorization evaluation is even more important than in standard settings.

stat.AP

A straightforward meta-analysis approach for oncology phase I dose-finding studies

Phase I early-phase clinical studies aim at investigating the safety and the underlying dose-toxicity relationship of a drug or combination. While little may still be known about the compound's properties, it is crucial to consider quantitative information available from any studies that may have been conducted previously on the same drug. A meta-analytic approach has the advantages of being able to properly account for between-study heterogeneity, and it may be readily extended to prediction or shrinkage applications. Here we propose a simple and robust two-stage approach for the estimation of maximum tolerated dose(s) (MTDs) utilizing penalized logistic regression and Bayesian random-effects meta-analysis methodology. Implementation is facilitated using standard R packages. The properties of the proposed methods are investigated in Monte-Carlo simulations. The investigations are motivated and illustrated by two examples from oncology.

stat.ME

Personalized Dynamic Treatment Regimes in Continuous Time: A Bayesian Approach for Optimizing Clinical Decisions with Timing

Accurate models of clinical actions and their impacts on disease progression are critical for estimating personalized optimal dynamic treatment regimes (DTRs) in medical/health research, especially in managing chronic conditions. Traditional statistical methods for DTRs usually focus on estimating the optimal treatment or dosage at each given medical intervention, but overlook the important question of "when this intervention should happen." We fill this gap by developing a two-step Bayesian approach to optimize clinical decisions with timing. In the first step, we build a generative model for a sequence of medical interventions-which are discrete events in continuous time-with a marked temporal point process (MTPP) where the mark is the assigned treatment or dosage. Then this clinical action model is embedded into a Bayesian joint framework where the other components model clinical observations including longitudinal medical measurements and time-to-event data conditional on treatment histories. In the second step, we propose a policy gradient method to learn the personalized optimal clinical decision that maximizes the patient survival by interacting the MTPP with the model on clinical observations while accounting for uncertainties in clinical observations learned from the posterior inference of the Bayesian joint model in the first step. A signature application of the proposed approach is to schedule follow-up visitations and assign a dosage at each visitation for patients after kidney transplantation. We evaluate our approach with comparison to alternative methods on both simulated and real-world datasets. In our experiments, the personalized decisions made by the proposed method are clinically useful: they are interpretable and successfully help improve patient survival.

stat.ME

Bayesian dose-regimen assessment in early phase oncology incorporating pharmacokinetics and pharmacodynamics

Phase I dose-finding trials in oncology seek to find the maximum tolerated dose (MTD) of a drug under a specific schedule. Evaluating drug-schedules aims at improving treatment safety while maintaining efficacy. However, while we can reasonably assume that toxicity increases with the dose for cytotoxic drugs, the relationship between toxicity and multiple schedules remains elusive. We proposed a Bayesian dose-regimen assessment method (DRtox) using pharmacokinetics/pharmacodynamics (PK/PD) information to estimate the maximum tolerated dose-regimen (MTD-regimen), at the end of the dose-escalation stage of a trial to be recommended for the next phase. We modeled the binary toxicity via a PD endpoint and estimated the dose-regimen toxicity relationship through the integration of a dose-regimen PD model and a PD toxicity model. For the dose-regimen PD model, we considered nonlinear mixed-effects models, and for the PD toxicity model, we proposed the following two Bayesian approaches: a logistic model and a hierarchical model. We evaluated the operating characteristics of the DRtox through simulation studies under various scenarios. The results showed that our method outperforms traditional model-based designs demonstrating a higher percentage of correctly selecting the MTD-regimen. Moreover, the inclusion of PK/PD information in the DRtox helped provide more precise estimates for the entire dose-regimen toxicity curve; therefore the DRtox may recommend alternative untested regimens for expansion cohorts. The DRtox should be applied at the end of the dose-escalation stage of an ongoing trial for patients with relapsed or refractory acute myeloid leukemia (NCT03594955) once all toxicity and PK/PD data are collected.

stat.ME

Clinical trials impacted by the COVID-19 pandemic: Adaptive designs to the rescue?

Very recently the new pathogen severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) was identified and the coronavirus disease 2019 (COVID-19) declared a pandemic by the World Health Organization. The pandemic has a number of consequences for the ongoing clinical trials in non-COVID-19 conditions. Motivated by four currently ongoing clinical trials in a variety of disease areas we illustrate the challenges faced by the pandemic and sketch out possible solutions including adaptive designs. Guidance is provided on (i) where blinded adaptations can help; (ii) how to achieve type I error rate control, if required; (iii) how to deal with potential treatment effect heterogeneity; (iv) how to utilize early readouts; and (v) how to utilize Bayesian techniques. In more detail approaches to resizing a trial affected by the pandemic are developed including considerations to stop a trial early, the use of group-sequential designs or sample size adjustment. All methods considered are implemented in a freely available R shiny app. Furthermore, regulatory and operational issues including the role of data monitoring committees are discussed.

stat.AP

Efficient adaptive designs for clinical trials of interventions for COVID-19

The COVID-19 pandemic has led to an unprecedented response in terms of clinical research activity. An important part of this research has been focused on randomized controlled clinical trials to evaluate potential therapies for COVID-19. The results from this research need to be obtained as rapidly as possible. This presents a number of challenges associated with considerable uncertainty over the natural history of the disease and the number and characteristics of patients affected, and the emergence of new potential therapies. These challenges make adaptive designs for clinical trials a particularly attractive option. Such designs allow a trial to be modified on the basis of interim analysis data or stopped as soon as sufficiently strong evidence has been observed to answer the research question, without compromising the trial's scientific validity or integrity. In this paper we describe some of the adaptive design approaches that are available and discuss particular issues and challenges associated with their use in the pandemic setting. Our discussion is illustrated by details of four ongoing COVID-19 trials that have used adaptive designs.

q-bio.QM

Random-effects meta-analysis of phase I dose-finding studies using stochastic process priors

Phase I dose-finding studies aim at identifying the maximal tolerated dose (MTD). It is not uncommon that several dose-finding studies are conducted, although often with some variation in the administration mode or dose panel. For instance, sorafenib (BAY 43-900) was used as monotherapy in at least 29 phase I trials according to a recent search in clinicaltrials.gov. Since the toxicity may not be directly related to the specific indication, synthesizing the information from several studies might be worthwhile. However, this is rarely done in practice and only a fixed-effect meta-analysis framework was proposed to date. We developed a Bayesian random-effects meta-analysis methodology to pool several phase I trials and suggest the MTD. A curve free hierarchical model on the logistic scale with random effects, accounting for between-trial heterogeneity, is used to model the probability of toxicity across the investigated doses. An Ornstein-Uhlenbeck Gaussian process is adopted for the random effects structure. Prior distributions for the curve free model are based on a latent Gamma process. An extensive simulation study showed good performance of the proposed method also under model deviations. Sharing information between phase I studies can improve the precision of MTD selection, at least when the number of trials is reasonably large.

q-bio.QM

Systematic reviews in paediatric multiple sclerosis and Creutzfeldt-Jakob disease exemplify shortcomings in methods used to evaluate therapies in rare conditions

BACKGROUND: Randomized controlled trials (RCTs) are the gold standard design of clinical research to assess interventions. However, RCTs cannot always be applied for practical or ethical reasons. To investigate the current practices in rare diseases, we review evaluations of therapeutic interventions in paediatric multiple sclerosis (MS) and Creutzfeldt-Jakob disease (CJD). In particular, we shed light on the endpoints used, the study designs implemented and the statistical methodologies applied. METHODS: We conducted literature searches to identify relevant primary studies. Data on study design, objectives, endpoints, patient characteristics, randomization and masking, type of intervention, control, withdrawals and statistical methodology were extracted from the selected studies. The risk of bias and the quality of the studies were assessed. RESULTS: Twelve (seven) primary studies on paediatric MS (CJD) were included in the qualitative synthesis. No double-blind, randomized placebo-controlled trial for evaluating interventions in paediatric MS has been published yet. Evidence from one open-label RCT is available. The observational studies are before-after studies or controlled studies. Three of the seven selected studies on CJD are RCTs, of which two received the maximum mark on the Oxford Quality Scale. Four trials are controlled observational studies. CONCLUSIONS: Evidence from double-blind RCTs on the efficacy of treatments appears to be variable between rare diseases. With regard to paediatric conditions it remains to be seen what impact regulators will have through e.g., paediatric investigation plans. Overall, there is space for improvement by using innovative trial designs and data analysis techniques.

q-bio.QM