SearcharxivSearch

arXiv subjects

Moritz Fabian Danzer

Publications and source records attributed to Moritz Fabian Danzer.

8 recordsLinked to original sources

Correcting for sampling variability in maximum likelihood-based one-sample log-rank tests

Single-arm studies in the early development phases of new treatments are not uncommon in the context of rare diseases or in paediatrics. If an assessment of efficacy is to be made at the end of such a study, the observed endpoints can be compared with reference values that can be derived from historical data. For a time-to-event endpoint, a statistical comparison with a reference curve can be made using the one-sample log-rank test. In order to ensure the interpretability of the results of this test, the role of the reference curve is crucial. This quantity is often estimated from a historical control group using a parametric procedure. Hence, it should be noted that it is subject to estimation uncertainty. However, this aspect is not taken into account in the one-sample log-rank test statistic. We analyse this estimation uncertainty for the common situation that the reference curve is estimated parametrically using the maximum likelihood method, and indicate how the variance estimation of the one-sample log-rank test can be adapted in order to take this variability into account. The resulting test procedures are illustrated using a data example and analysed in more detail using simulations, particularly in comparison with established two-sample methods.

stat.ME

Exhausting the type I error level in event-driven group-sequential designs with a closed testing procedure for progression-free and overall survival

In oncological clinical trials, overall survival (OS) is the gold-standard endpoint, but long follow-up and treatment switching can delay or dilute detectable effects. Progression-free survival (PFS) often provides earlier evidence and is therefore frequently used together with OS as multiple primary endpoints. Since in certain scenarios trial success may be defined if one of the two hypotheses involved can be rejected, a correction for multiple testing may be deemed necessary. Because PFS and OS are generally highly dependent, their test statistics are typically correlated. Ignoring this dependency (e.g. via a simple Bonferroni correction) is not power optimal. We develop a group-sequential testing procedure for the multiple primary endpoints PFS and OS that fully exhausts the family-wise error rate (FWER) by exploiting their dependence. Specifically, we characterize the joint asymptotic distribution of log-rank statistics across endpoints and multiple event-driven analysis cutoffs. Furthermore, we show that we can consistently estimate the covariance structure. Embedding these results in a closed testing procedure, we can recalculate critical values of the test statistics in order to spend the available type I error optimally. An important extension to the current literature is that we allow for both interim and final analysis to be event-driven. Simulations based on illness-death multi-state models empirically confirm FWER control for moderate to large sample sizes. Compared with a simple Bonferroni correction, the proposed methods recover roughly two thirds of the power loss for OS, increase disjunctive and conjunctive power, and enable meaningful early stopping. In planning, these gains translate into about 5% fewer OS events required to reach the targeted power. We also discuss practical issues in the implementation of such designs and possible extensions of the introduced method.

stat.ME

Confirmatory Adaptive Hypothesis Tests in Markovian Illness-Death Models

Classic adaptive designs for time-to-event trials are based on the log-rank statistic and its increments. Thereby, only information from the time-to-event endpoint on which the selected log-rank statistic is based may be used for data-dependent design modifications in interim analyses. Further information (e.g. surrogate parameters) may not be used. As pointed out in a letter by P. Bauer and M. Posch in 2004, adaptive tests on overall survival (OS) based on the log-rank statistic do in general not control the significance level if interim information on progression-free survival (PFS) is used for sample size adjustments, because progression is associated with increased risk of death. In contrast, in adaptive designs for time-to-event trials, which are constructed according to the principle of patient-wise separation, all trial data observed in interim analyses may be used for design modifications without compromizing type one error rate control. But by design, this comes at the price of incomplete use of the primary endpoint data in the final test decision or worst-case considerations which lead to a loss of power. Thus, the patient-wise separation approach cannot be regarded as a general solution to the problem described by Bauer and Posch. We address this problem within the framework of a comprehensive independent increments approach. We develop adaptive tests on OS in which sample size adjustments may be based on the observed interim data of both OS and PFS, while avoiding the problems of the patient-wise separation approach. We provide this methodology for both single-arm trials, in which a new therapy is compared with a pre-specified deterministic reference, and randomized trials, in which a new therapy is compared with a concurrent control group. The underlying assumption is that the joint distribution of OS and PFS is induced by a Markovian illness-death model.

stat.ME

Adaptive weight selection for time-to-event data under non-proportional hazards

When planning a clinical trial for a time-to-event endpoint, we require an estimated effect size and need to consider the type of effect. Usually, an effect of proportional hazards is assumed with the hazard ratio as the corresponding effect measure. Thus, the standard procedure for survival data is generally based on a single-stage log-rank test. Knowing that the assumption of proportional hazards is often violated and sufficient knowledge to derive reasonable effect sizes is usually unavailable, such an approach is relatively rigid. We introduce a more flexible procedure by combining two methods designed to be more robust in case we have little to no prior knowledge. First, we employ a more flexible adaptive multi-stage design instead of a single-stage design. Second, we apply combination-type tests in the first stage of our suggested procedure to benefit from their robustness under uncertainty about the deviation pattern. We can then use the data collected during this period to choose a more specific single-weighted log-rank test for the subsequent stages. In this step, we employ Royston-Parmar spline models to extrapolate the survival curves to make a reasonable decision. Based on a real-world data example, we show that our approach can save a trial that would otherwise end with an inconclusive result. Additionally, our simulation studies demonstrate a sufficient power performance while maintaining more flexibility.

stat.ME

Confirmatory adaptive group sequential designs for clinical trials with multiple time-to-event outcomes in Markov models

The analysis of multiple time-to-event outcomes in a randomised controlled clinical trial can be accomplished with exisiting methods. However, depending on the characteristics of the disease under investigation and the circumstances in which the study is planned, it may be of interest to conduct interim analyses and adapt the study design if necessary. Due to the expected dependency of the endpoints, the full available information on the involved endpoints may not be used for this purpose. We suggest a solution to this problem by embedding the endpoints in a multi-state model. If this model is Markovian, it is possible to take the disease history of the patients into account and allow for data-dependent design adaptiations. To this end, we introduce a flexible test procedure for a variety of applications, but are particularly concerned with the simultaneous consideration of progression-free survival (PFS) and overall survival (OS). This setting is of key interest in oncological trials. We conduct simulation studies to determine the properties for small sample sizes and demonstrate an application based on data from the NB2004-HR study.

stat.ME

A statistical framework for planning and analysing test-retest studies for repeatability of quantitative biomarker measurements

There is an increasing number of potential biomarkers that could allow for early assessment of treatment response or disease progression. However, measurements of quantitative biomarkers are subject to random variability. Hence, differences of a biomarker in longitudinal measurements do not necessarily represent real change but might be caused by this random measurement variability. Before utilizing a quantitative biomarker in longitudinal studies, it is therefore essential to assess the measurement repeatability. Measurement repeatability obtained from test-retest studies can be quantified by the repeatability coefficient (RC), which is then used in the subsequent longitudinal study to determine if a measured difference represents real change or is within the range of expected random measurement variability. The quality of the point estimate of RC therefore directly governs the assessment quality of the longitudinal study. RC estimation accuracy depends on the case number in the test-retest study, but despite its pivotal role, no comprehensive framework for sample size calculation of test-retest studies exists. To address this issue, we have established such a framework, which allows for flexible sample size calculation of test-retest studies, based upon newly introduced criteria concerning assessment quality in the longitudinal study. This also permits retrospective assessment of prior test-retest studies.

stat.ME

On variance estimation for the one-sample log-rank test

Time-to-event endpoints show an increasing popularity in phase II cancer trials. The standard statistical tool for such one-armed survival trials is the one-sample log-rank test. Its distributional properties are commonly derived in the large sample limit. It is however known from the literature, that the asymptotical approximations suffer when sample size is small. There have already been several attempts to address this problem. While some approaches do not allow easy power and sample size calculations, others lack a clear theoretical motivation and require further considerations. The problem itself can partly be attributed to the dependence of the compensated counting process and its variance estimator. For this purpose, we suggest a variance estimator which is uncorrelated to the compensated counting process. Moreover, this and other present approaches to variance estimation are covered as special cases by our general framework. For practical application, we provide sample size and power calculations for any approach fitting into this framework. Finally, we use simulations and real world data to study the empirical type I error and power performance of our methodology as compared to standard approaches.

stat.ME

One-sample log-rank tests with consideration of reference curve sampling variability

The one-sample log-rank test is the method of choice for single-arm Phase II trials with time-to-event endpoint. It allows to compare the survival of the patients to a reference survival curve that typically represents the expected survival under standard of care. The classical one-sample log-rank test, however, assumes that the reference survival curve is deterministic. This ignores that the reference curve is commonly estimated from historic data and thus prone to statistical error. Ignoring sampling variability of the reference curve results in type I error rate inflation. For that reason, a new one-sample log-rank test is proposed that explicitly accounts for the statistical error made in the process of estimating the reference survival curve. The test statistic and its distributional properties are derived using martingale techniques in the large sample limit. In particular, a sample size formula is provided. Small sample properties regarding type I and type II error rate control are studied by simulation. A case study is conducted to study the influence of several design parameters of a single-armed trial on the inflation of the type I error rate when reference curve sampling variability is ignored.

stat.ME