SearcharxivSearch

arXiv subjects

Tim Friede

Publications and source records attributed to Tim Friede.

At least 37 records · Page 2Linked to original sources

A studentized permutation test in group sequential designs

In group sequential designs, where several data looks are conducted for early stopping, we generally assume the vector of test statistics from the sequential analyses follows (at least approximately or asymptotially) a multivariate normal distribution. However, it is well-known that test statistics for which an asymptotic distribution is derived may suffer from poor small sample approximation. This might become even worse with an increasing number of data looks. The aim of this paper is to improve the small sample behaviour of group sequential designs while maintaining the same asymptotic properties as classical group sequential designs. This improvement is achieved through the application of a modified permutation test. In particular, this paper shows that the permutation distribution approximates the distribution of the test statistics not only under the null hypothesis but also under the alternative hypothesis, resulting in an asymptotically valid permutation test. An extensive simulation study shows that the proposed permutation test better controls the Type I error rate than its competitors in the case of small sample sizes.

math.ST

Survival analysis for AdVerse events with VarYing follow-up times (SAVVY): summary of findings and a roadmap for the future of safety analyses in clinical trials

The SAVVY project aims to improve the analyses of adverse events (AEs) in clinical trials through the use of survival techniques appropriately dealing with varying follow-up times and competing events (CEs). This paper summarizes key features and conclusions from the various SAVVY papers. Through theoretical investigations using simulations and in an empirical study including randomized clinical trials from several sponsor organisations, biases from ignoring varying follow-up times or CEs are investigated. The bias of commonly used estimators of the absolute and relative AE risk is quantified. Furthermore, we provide a cursory assessment of how pertinent guidelines for the analysis of safety data deal with the features of varying follow-up time and CEs. SAVVY finds that for both, avoiding bias and categorization of evidence with respect to treatment effect on AE risk into categories, the choice of the estimator is key and more important than features of the underlying data such as percentage of censoring, CEs, amount of follow-up, or value of the gold-standard. The choice of the estimator of the cumulative AE probability and the definition of CEs are crucial. SAVVY recommends using the Aalen-Johansen estimator (AJE) with an appropriate definition of CEs whenever the risk for AEs is to be quantified. There is an urgent need to improve the guidelines of reporting AEs so that incidence proportions or one minus Kaplan-Meier estimators are finally replaced by the AJE with appropriate definition of CEs.

stat.AP

A two-step approach for analyzing time to event data under non-proportional hazards

The log-rank test and the Cox proportional hazards model are commonly used to compare time-to-event data in clinical trials, as they are most powerful under proportional hazards. But there is a loss of power if this assumption is violated, which is the case for some new oncology drugs like immunotherapies. We consider a two-stage test procedure, in which the weighting of the log-rank test statistic depends on a pre-test of the proportional hazards assumption. I.e., depending on the pre-test either the log-rank or an alternative test is used to compare the survival probabilities. We show that if naively implemented this can lead to a substantial inflation of the type-I error rate. To address this, we embed the two-stage test in a permutation test framework to keep the nominal level alpha. We compare the operating characteristics of the two-stage test with the log-rank test and other tests by clinical trial simulations.

stat.ME

Methods for non-proportional hazards in clinical trials: A systematic review

For the analysis of time-to-event data, frequently used methods such as the log-rank test or the Cox proportional hazards model are based on the proportional hazards assumption, which is often debatable. Although a wide range of parametric and non-parametric methods for non-proportional hazards (NPH) has been proposed, there is no consensus on the best approaches. To close this gap, we conducted a systematic literature search to identify statistical methods and software appropriate under NPH. Our literature search identified 907 abstracts, out of which we included 211 articles, mostly methodological ones. Review articles and applications were less frequently identified. The articles discuss effect measures, effect estimation and regression approaches, hypothesis tests, and sample size calculation approaches, which are often tailored to specific NPH situations. Using a unified notation, we provide an overview of methods available. Furthermore, we derive some guidance from the identified articles. We summarized the contents from the literature review in a concise way in the main text and provide more detailed explanations in the supplement.

stat.ME

How trace plots help interpret meta-analysis results

The trace plot is seldom used in meta-analysis, yet it is a very informative plot. In this article we define and illustrate what the trace plot is, and discuss why it is important. The Bayesian version of the plot combines the posterior density of tau, the between-study standard deviation, and the shrunken estimates of the study effects as a function of tau. With a small or moderate number of studies, tau is not estimated with much precision, and parameter estimates and shrunken study effect estimates can vary widely depending on the correct value of tau. The trace plot allows visualization of the sensitivity to tau along with a plot that shows which values of tau are plausible and which are implausible. A comparable frequentist or empirical Bayes version provides similar results. The concepts are illustrated using examples in meta-analysis and meta-regression; implementaton in R is facilitated in a Bayesian or frequentist framework using the bayesmeta and metafor packages, respectively.

stat.ME

Subgroup identification using individual participant data from multiple trials on low back pain

Model-based recursive partitioning (MOB) and its extension, metaMOB, are potent tools for identifying subgroups with differential treatment effects. In the metaMOB approach random effects are used to model heterogeneity of the treatment effects when pooling data from various trials. In situations where interventions offer only small overall benefits and require extensive, costly trials with a large participant enrollment, leveraging individual-participant data (IPD) from multiple trials can help identify individuals who are most likely to benefit from the intervention. We explore the application of MOB and metaMOB in the context of non specific low back pain treatment, using synthesized data based on a subset of the individual participant data meta-analysis by Patel et al. Our study underscores the need to explore heterogeneity in intercepts and treatment effects to identify subgroups with differential treatment effects in IPD meta-analyses.

stat.ME

A neutral comparison of statistical methods for time-to-event analyses under non-proportional hazards

While well-established methods for time-to-event data are available when the proportional hazards assumption holds, there is no consensus on the best inferential approach under non-proportional hazards (NPH). However, a wide range of parametric and non-parametric methods for testing and estimation in this scenario have been proposed. To provide recommendations on the statistical analysis of clinical trials where non proportional hazards are expected, we conducted a comprehensive simulation study under different scenarios of non-proportional hazards, including delayed onset of treatment effect, crossing hazard curves, subgroups with different treatment effect and changing hazards after disease progression. We assessed type I error rate control, power and confidence interval coverage, where applicable, for a wide range of methods including weighted log-rank tests, the MaxCombo test, summary measures such as the restricted mean survival time (RMST), average hazard ratios, and milestone survival probabilities as well as accelerated failure time regression models. We found a trade-off between interpretability and power when choosing an analysis strategy under NPH scenarios. While analysis methods based on weighted logrank tests typically were favorable in terms of power, they do not provide an easily interpretable treatment effect estimate. Also, depending on the weight function, they test a narrow null hypothesis of equal hazard functions and rejection of this null hypothesis may not allow for a direct conclusion of treatment benefit in terms of the survival function. In contrast, non-parametric procedures based on well interpretable measures as the RMST difference had lower power in most scenarios. Model based methods based on specific survival distributions had larger power, however often gave biased estimates and lower than nominal confidence interval coverage.

stat.ME

Robust Confidence Intervals for Meta-Regression with Interaction Effects

Meta-analysis is an important statistical technique for synthesizing the results of multiple studies regarding the same or closely related research question. So-called meta-regression extends meta-analysis models by accounting for studylevel covariates. Mixed-effects meta-regression models provide a powerful tool for evidence synthesis, by appropriately accounting for betweem-study heterogeneity. In fact, modelling the study effect in terms of random effects and moderators not only allows to examine the impact of the moderators, but often leads to more accurate estimates of the involved parameters. Nevertheless, due to the often small number of studies on a specific research topic, interactions are often neglected in meta-regression. In this work, we consider the research questions (i) how moderator interactions influence inference in mixed-effects meta-regression models and (ii) whether some inference methods are more reliable than others. Here, we review robust methods for confidence intervals in meta-regression models including interaction effects. These methods are based on the application of robust sandwich estimators for estimating the variance-covariance matrix of the vector of model coefficients. Furthermore, we compare different versions of these robust estimators in an extensive simulation study. We thereby investigate coverage and length of seven different confidence intervals under varying conditions. We conclude with some practical recommendations.

stat.ME

The impact of neglected confounding and interactions in mixed-effects meta-regression

Analysts seldom include interaction terms in meta-regression model, what can introduce bias if an interaction is present. We illustrate this in the current paper by re-analyzing an example from research on acute heart failure, where neglecting an interaction might have led to erroneous inference and conclusions. Moreover, we perform a brief simulation study based on this example highlighting the effects caused by omitting or unnecessarily including interaction terms. Based on our results, we recommend to always include interaction terms in mixed-effects meta-regression models, when such interactions are plausible.

stat.ME

On the role of benchmarking data sets and simulations in method comparison studies

Method comparisons are essential to provide recommendations and guidance for applied researchers, who often have to choose from a plethora of available approaches. While many comparisons exist in the literature, these are often not neutral but favour a novel method. Apart from the choice of design and a proper reporting of the findings, there are different approaches concerning the underlying data for such method comparison studies. Most manuscripts on statistical methodology rely on simulation studies and provide a single real-world data set as an example to motivate and illustrate the methodology investigated. In the context of supervised learning, in contrast, methods are often evaluated using so-called benchmarking data sets, i.e. real-world data that serve as gold standard in the community. Simulation studies, on the other hand, are much less common in this context. The aim of this paper is to investigate differences and similarities between these approaches, to discuss their advantages and disadvantages and ultimately to develop new approaches to the evaluation of methods picking the best of both worlds. To this aim, we borrow ideas from different contexts such as mixed methods research and Clinical Scenario Evaluation.

stat.ME

Accounting for Time Dependency in Meta-Analyses of Concordance Probability Estimates

Recent years have seen the development of many novel scoring tools for disease prognosis and prediction. To become accepted for use in clinical applications, these tools have to be validated on external data. In practice, validation is often hampered by logistical issues, resulting in multiple small-sized validation studies. It is therefore necessary to synthesize the results of these studies using techniques for meta-analysis. Here we consider strategies for meta-analyzing the concordance probability for time-to-event data ("C-index"), which has become a popular tool to evaluate the discriminatory power of prediction models with a right-censored outcome. We show that standard meta-analysis of the C-index may lead to biased results, as the magnitude of the concordance probability depends on the length of the time interval used for evaluation (defined e.g. by the follow-up time, which might differ considerably between studies). To address this issue, we propose a set of methods for random-effects meta-regression that incorporate time directly as covariate in the model equation. In addition to analyzing nonlinear time trends via fractional polynomial, spline, and exponential decay models, we provide recommendations on suitable transformations of the C-index before meta-regression. Our results suggest that the C-index is best meta-analyzed using fractional polynomial meta-regression with logit-transformed C-index values. Classical random-effects meta-analysis (not considering time as covariate) is demonstrated to be a suitable alternative when follow-up times are small. Our findings have implications for the reporting of C-index values in future studies, which should include information on the length of the time interval underlying the calculations.

stat.ME

Using the bayesmeta R package for Bayesian random-effects meta-regression

BACKGROUND: Random-effects meta-analysis within a hierarchical normal modeling framework is commonly implemented in a wide range of evidence synthesis applications. More general problems may even be tackled when considering meta-regression approaches that in addition allow for the inclusion of study-level covariables. METHODS: We describe the Bayesian meta-regression implementation provided in the bayesmeta R package including the choice of priors, and we illustrate its practical use. RESULTS: A wide range of example applications are given, such as binary and continuous covariables, subgroup analysis, indirect comparisons, and model selection. Example R code is provided. CONCLUSIONS: The bayesmeta package provides a flexible implementation. Due to the avoidance of MCMC methods, computations are fast and reproducible, facilitating quick sensitivity checks or large-scale simulation studies.

stat.CO

Summarizing empirical information on between-study heterogeneity for Bayesian random-effects meta-analysis

In Bayesian meta-analysis, the specification of prior probabilities for the between-study heterogeneity is commonly required, and is of particular benefit in situations where only few studies are included. Among the considerations in the set-up of such prior distributions, the consultation of available empirical data on a set of relevant past analyses sometimes plays a role. How exactly to summarize historical data sensibly is not immediately obvious; in particular, the investigation of an empirical collection of heterogeneity estimates will not target the actual problem and will usually only be of limited use. The commonly used normal-normal hierarchical model for random-effects meta-analysis is extended to infer a heterogeneity prior. Using an example data set, we demonstrate how to fit a distribution to empirically observed heterogeneity data from a set of meta-analyses. Considerations also include the choice of a parametric distribution family. Here, we focus on simple and readily applicable approaches to then translate these into (prior) probability distributions.

stat.ME

Survival analysis for AdVerse events with VarYing follow-up times (SAVVY) -- comparison of adverse event risks in randomized controlled trials

Analyses of adverse events (AEs) are an important aspect of the evaluation of experimental therapies. The SAVVY (Survival analysis for AdVerse events with Varying follow-up times) project aims to improve the analyses of AE data in clinical trials through the use of survival techniques appropriately dealing with varying follow-up times, censoring, and competing events (CE). In an empirical study including seventeen randomized clinical trials the effect of varying follow-up times, censoring, and competing events on comparisons of two treatment arms with respect to AE risks is investigated. The comparisons of relative risks (RR) of standard probability-based estimators to the gold-standard Aalen-Johansen estimator or hazard-based estimators to an estimated hazard ratio (HR) from Cox regression are done descriptively, with graphical displays, and using a random effects meta-analysis on AE level. The influence of different factors on the size of the bias is investigated in a meta-regression. We find that for both, avoiding bias and categorization of evidence with respect to treatment effect on AE risk into categories, the choice of the estimator is key and more important than features of the underlying data such as percentage of censoring, CEs, amount of follow-up, or value of the gold-standard RR. There is an urgent need to improve the guidelines of reporting AEs so that incidence proportions are finally replaced by the Aalen-Johansen estimator - rather than by Kaplan-Meier - with appropriate definition of CEs. For RRs based on hazards, the HR based on Cox regression has better properties than the ratio of incidence densities.

stat.AP

Model-based recursive partitioning for discrete event times

Model-based recursive partitioning (MOB) is a semi-parametric statistical approach allowing the identification of subgroups that can be combined with a broad range of outcome measures including continuous time-to-event outcomes. When time is measured on a discrete scale, methods and models need to account for this discreetness as otherwise subgroups might be spurious and effects biased. The test underlying the splitting criterion of MOB, the M-fluctuation test, assumes independent observations. However, for fitting discrete time-to-event models the data matrix has to be modified resulting in an augmented data matrix violating the independence assumption. We propose MOB for discrete Survival data (MOB-dS) which controls the type I error rate of the test used for data splitting and therefore the rate of identifying subgroups although none is present. MOB-ds uses a permutation approach accounting for dependencies in the augmented time-to-event data to obtain the distribution under the null hypothesis of no subgroups being present. Through simulations we investigate the type I error rate of the new MOB-dS and the standard MOB for different patterns of survival curves and event rates. We find that the type I error rates of the test is well controlled for MOB-dS, but observe some considerable inflations of the error rate for MOB. To illustrate the proposed methods, MOB-dS is applied to data on unemployment duration.

stat.ME

Double arcsine transform not appropriate for meta-analysis

The variance-stabilizing Freeman-Tukey double arcsine transform was originally proposed for inference on single proportions. Subsequently, its use has been suggested in the context of meta-analysis of proportions. While some erratic behaviour has been observed previously, here we point out and illustrate general issues of monotonicity and invertibility that make this transform unsuitable for meta-analysis purposes.

stat.ME

Coping with Information Loss and the Use of Auxiliary Sources of Data: A Report from the NISS Ingram Olkin Forum Series on Unplanned Clinical Trial Disruptions

Clinical trials disruption has always represented a non negligible part of the ending of interventional studies. While the SARS-CoV-2 (COVID-19) pandemic has led to an impressive and unprecedented initiation of clinical research, it has also led to considerable disruption of clinical trials in other disease areas, with around 80% of non-COVID-19 trials stopped or interrupted during the pandemic. In many cases the disrupted trials will not have the planned statistical power necessary to yield interpretable results. This paper describes methods to compensate for the information loss arising from trial disruptions by incorporating additional information available from auxiliary data sources. The methods described include the use of auxiliary data on baseline and early outcome data available from the trial itself and frequentist and Bayesian approaches for the incorporation of information from external data sources. The methods are illustrated by application to the analysis of artificial data based on the Primary care pediatrics Learning Activity Nutrition (PLAN) study, a clinical trial assessing a diet and exercise intervention for overweight children, that was affected by the COVID-19 pandemic. We show how all of the methods proposed lead to an increase in precision relative to use of complete case data only.

stat.AP

On the role of data, statistics and decisions in a pandemic

A pandemic poses particular challenges to decision-making because of the need to continuously adapt decisions to rapidly changing evidence and available data. For example, which countermeasures are appropriate at a particular stage of the pandemic? How can the severity of the pandemic be measured? What is the effect of vaccination in the population and which groups should be vaccinated first? The process of decision-making starts with data collection and modeling and continues to the dissemination of results and the subsequent decisions taken. The goal of this paper is to give an overview of this process and to provide recommendations for the different steps from a statistical perspective. In particular, we discuss a range of modeling techniques including mathematical, statistical and decision-analytic models along with their applications in the COVID-19 context. With this overview, we aim to foster the understanding of the goals of these modeling approaches and the specific data requirements that are essential for the interpretation of results and for successful interdisciplinary collaborations. A special focus is on the role played by data in these different models, and we incorporate into the discussion the importance of statistical literacy, and of effective dissemination and communication of findings.

stat.OT