SearcharxivSearch

arXiv subjects

Leonhard Held

Publications and source records attributed to Leonhard Held.

At least 19 recordsLinked to original sources

Edgington's Combination Method for Two-Study Meta-Analysis: An Empirical Evaluation in 1226 Meta-Analyses

Two-study meta-analyses are common in evidence synthesis but pose major statistical challenges. With only two studies, the between-study variance cannot be reliably estimated, rendering standard random-effects methods unstable. Here, we investigate meta-analyses based on Edgington's p-value combination method as an alternative approach, applying it to 1226 two-study meta-analyses from the German Institute for Quality and Efficiency in Health Care (IQWiG). Like fixed-effect meta-analysis, Edgington's method is calibrated under homogeneity. However, it adapts confidence interval width to observed between-study discrepancy without requiring explicit heterogeneity estimation. In all of the examined meta-analyses, this leads to confidence intervals that contain both study-specific estimates but remain informative. Edgington's method agrees with fixed-effect meta-analysis on statistical significance (at two-sided $\alpha$ = 0.05) in 91% of all meta-analyses, but can give wider intervals when study results are discrepant and narrower intervals when results are highly consistent. Weighted extensions of Edgington's method shift point estimates toward the more precise study while preserving much of this adaptive behavior. We conclude that Edgington's method offers a principled and practically useful complement to existing approaches for two-study meta-analysis, occupying a middle ground between standard fixed-effect and random-effects approaches.

stat.AP

Adjusting for Outcome Reporting Bias in Meta-analysis: A Multiple Imputation Approach

Background: Outcome reporting bias (ORB) occurs when study outcomes are selectively reported based on their results. ORB potentially undermines the credibility and validity of meta-analyses and contributes to research waste by distorting overall treatment effects. ORB can be viewed as a missing data problem in which unreported study outcomes introduce bias. Despite the serious implications ORB poses, it remains an underrecognized issue, with only a few adjustment methods available. Methods: We propose an approach that addresses unreported study outcomes in meta-analyses through multiple imputation for univariate and multivariate meta-analysis. To assess the impact of ORB in meta-analyses, we apply our proposed methodology to real clinical data affected by ORB, and conduct a simulation study to evaluate the method's performance under a range of scenarios. Results: The proposed method provides bias-adjusted estimates under assumed selective non-reporting mechanisms. In the application to clinical data, ORB-adjusted estimates were systematically shifted towards less extreme treatment effects compared with naive analyses, highlighting the potential magnitude of ORB in practice. The simulation study shows that the extent of adjustment depends on the assumed selection mechanism and the degree of heterogeneity, with stronger selection leading to larger adjustment. Conclusions: Imputing unreported study outcomes provides a promising approach to address ORB in meta-analyses. The multivariate approach extends ORB adjustment to jointly model correlated outcomes, allowing borrowing of strength across outcomes. Overall, we propose a practical and flexible approach for evaluating the sensitivity of univariate and multivariate meta-analytic conclusions to ORB.

stat.ME

Bayes Factor Group Sequential Designs

The Bayes factor, the data-based updating factor from prior to posterior odds, is a principled measure of relative evidence for two competing hypotheses. It is naturally suited to sequential data analysis in settings such as clinical trials and animal experiments, where early stopping for efficacy or futility is desirable. However, designing such studies is challenging because computing design characteristics, such as the probability of obtaining conclusive evidence or the expected sample size, typically requires computationally intensive Monte Carlo simulations, as no closed-form or efficient numerical methods exist. To address this issue, we extend results from classical group sequential design theory to sequential Bayes factor designs. The key idea is to derive Bayes factor stopping regions in terms of the z-statistic and use the known distribution of the cumulative z-statistics to compute stopping probabilities through multivariate normal integration. The resulting method is fast, accurate, and simulation-free. We illustrate it with examples from clinical trials, animal experiments, and psychological studies. We also provide an open-source implementation in the bfpwr R package. Our method makes exploring sequential Bayes factor designs as straightforward as classical group sequential designs, enabling experiments to rapidly design informative and efficient experiments.

stat.ME

Balancing Evidentiary Value and Sample Size of Adaptive Designs with Application to Animal Experiments

Reducing the number of experimental units is one of the three pillars of the 3R principles (Replace, Reduce, Refine) in animal research. At the same time, statistical error rates need to be controlled to enable reliable inferences and decisions. This paper proposes to adopt diagnostic likelihood ratios and the diagnostic odds ratio to statistical hypothesis tests and to adjust it for sample size to obtain a novel measure to quantify for the evidentiary value of one experimental unit. The experimental unit information index (EUII) is based on power, Type-I error and sample size, and has attractive interpretations both in terms of frequentist error rates and Bayesian posterior odds. We introduce the EUII in simple statistical test settings and show that its asymptotic value depends only on the assumed relative effect size under the alternative. We then extend the definition to adaptive designs where early stopping for efficacy or futility may cause reductions in sample size. Application to group-sequential designs show the usefulness of the approach when the goal is to maximize the evidentiary value of one experimental unit. A reanalysis of 2738 animal experiments with simulated results from (post-hoc) interim analyses illustrates the possible savings in sample size.

stat.ME

Prediction intervals for random-effects meta-analysis based on confidence distributions and Edgington's method

Statistical inference about the average effect in random-effects meta-analysis has been considered insufficient in the presence of substantial between-study heterogeneity. Predictive distributions are well-suited for quantifying heterogeneity since they are interpretable on the effect scale and provide clinically relevant information about future events. We construct predictive distributions accounting for uncertainty through confidence distributions from Edgington's $p$-value combination method and the generalized heterogeneity statistic. Simulation results suggest that 95% prediction intervals typically achieve nominal coverage when more than three studies are available and effectively reflect skewness of effect estimates in scenarios with 20 or less studies. Formulations that ignore uncertainty in heterogeneity estimation typically fail to achieve correct coverage, underscoring the need for this adjustment in random-effects meta-analysis.

stat.ME

Edgington's Method for Random-Effects Meta-Analysis Part I: Estimation

Meta-analysis can be formulated as combining $p$-values across studies into a joint $p$-value function, from which point estimates and confidence intervals can be derived. We extend the meta-analytic estimation framework based on combined $p$-value functions to incorporate uncertainty in heterogeneity estimation by employing a confidence distribution approach. Specifically, the confidence distribution of Edgington's method is adjusted according to the confidence distribution of the heterogeneity parameter constructed from the generalized heterogeneity statistic. Simulation results suggest that 95% confidence intervals approach nominal coverage under most scenarios involving more than three studies and heterogeneity. Under no heterogeneity or for only three studies, the confidence interval typically overcovers, but is often narrower than the Hartung-Knapp-Sidik-Jonkman interval. The point estimator exhibits small bias under model misspecification and moderate to large heterogeneity. Edgington's method provides a practical alternative to classical approaches, with adjustment for heterogeneity estimation uncertainty often improving confidence interval coverage.

stat.ME

Stabilizing Thompson Sampling with Null Hypothesis Bayesian Response-Adaptive Randomization

Response-adaptive randomization (RAR) methods can be used to adapt randomization probabilities based on accumulating data, aiming to increase the probability of allocating patients to effective treatments. A popular RAR method is Thompson sampling, which randomizes patients proportionally to the Bayesian posterior probability that each treatment is the most effective. However, its high variability can also increase the risk of assigning patients to inferior treatments and lead to inferential problems such as confidence interval undercoverage. We propose a principled method based on Bayesian hypothesis testing to address these issues: We introduce a null hypothesis postulating equal effectiveness of treatments. Bayesian model averaging then induces shrinkage toward equal randomization probabilities, with the degree of shrinkage controlled by the prior probability of the null hypothesis. Equal randomization and Thompson sampling arise as special cases when the prior probability is one or zero, respectively. A simulation study demonstrates that the method can mitigate issues with Thompson sampling and has comparable statistical properties to Thompson sampling with common ad hoc modifications such as power transformation and probability capping. Under the null hypothesis and a normal model, the randomization probabilities are shown to converge asymptotically to equal randomization, unlike those of Thompson sampling. We implement the method in the free and open-source R package brar, enabling experimenters to easily perform null hypothesis Bayesian RAR and support more effective randomization of patients.

stat.ME

Combined P-value Functions for Compatible Effect Estimation and Hypothesis Testing in Drug Regulation

The two-trials rule in drug regulation requires statistically significant results from two pivotal trials to demonstrate efficacy. However, it is unclear how the effect estimates from both trials should be combined to quantify the drug effect. Fixed-effect meta-analysis is commonly used but may yield confidence intervals that exclude the value of no effect even when the two-trials rule is not fulfilled. We systematically address this by recasting the two-trials rule and meta-analysis in a unified framework of combined p-value functions, where they are variants of Wilkinson's and Stouffer's combination methods, respectively. This allows us to obtain compatible combined p-values, effect estimates, and confidence intervals, which we derive in closed-form. Additionally, we provide new results for Edgington's, Fisher's, Pearson's, and Tippett's p-value combination methods. When both trials have the same true effect, all methods can consistently estimate it, although some show bias. When true effects differ, the two-trials rule and Pearson's method are conservative (converging to the less extreme effect), Fisher's and Tippett's methods are anti-conservative (converging to the more extreme effect), and Edgington's method and meta-analysis are balanced (converging to a weighted average). Notably, Edgington's confidence intervals asymptotically always include the individual trial effects, while meta-analytic confidence intervals shrink to a point at the weighted average effect. We conclude that all of these methods may be appropriate depending on the estimand of interest. We implement combined p-value function inference for two trials in the R package twotrials, allowing researchers to easily perform compatible hypothesis testing and effect estimation.

stat.ME

Assessing the replicability of RCTs in RWE emulations

Background: The standard regulatory approach to assess replication success is the two-trials rule, requiring both the original and the replication study to be significant with effect estimates in the same direction. The sceptical p-value was recently presented as an alternative method for the statistical assessment of the replicability of study results. Methods: We review the statistical properties of the sceptical p-value and compare those to the two-trials rule. We extend the methodology to non-inferiority trials and describe how to invert the sceptical p-value to obtain confidence intervals. We illustrate the performance of the different methods using real-world evidence emulations of randomized, controlled trials (RCTs) conducted within the RCT DUPLICATE initiative. Results: The sceptical p-value depends not only on the two p-values, but also on sample size and effect size of the two studies. It can be calibrated to have the same Type-I error rate as the two-trials rule, but has larger power to detect an existing effect. In the application to the results from the RCT DUPLICATE initiative, the sceptical p- value leads to qualitatively similar results than the two-trials rule, but tends to show more evidence for treatment effects compared to the two-trials rule. Conclusion: The sceptical p-value represents a valid statistical measure to assess the replicability of study results and is especially useful in the context of real-world evidence emulations.

stat.ME

A comparison of combined p-value functions for meta-analysis

P-value functions are modern statistical tools that unify effect estimation and hypothesis testing and can provide alternative point and interval estimates compared to standard meta-analysis methods, using any of the many $p$-value combination procedures available (Xie et al., 2011, JASA). We provide a systematic comparison of different combination procedures, both from a theoretical perspective and through simulation. We show that many prominent p-value combination methods (e.g. Fisher's method) are not invariant to the orientation of the underlying one-sided p-values. Only Edgington's method, a lesser-known combination method based on the sum of $p$-values, is orientation-invariant and still provides confidence intervals not restricted to be symmetric around the point estimate. Adjustments for heterogeneity can also be made and results from a simulation study indicate that Edgington's method can compete with more standard meta-analytic methods.

stat.ME

Addressing Outcome Reporting Bias in Meta-analysis: A Selection Model Perspective

Outcome Reporting Bias (ORB) poses significant threats to the validity of meta-analytic findings. It occurs when researchers selectively report outcomes based on the significance or direction of results, potentially leading to distorted treatment effect estimates. Despite its critical implications, ORB remains an under-recognized issue, with few comprehensive adjustment methods available. The goal of this research is to investigate ORB-adjustment techniques through a selection model lens, thereby extending some of the existing methodological approaches available in the literature. To gain a better insight into the effects of ORB in meta-analysis of clinical trials, specifically in the presence of heterogeneity, and to assess the effectiveness of ORB-adjustment techniques, we apply the methodology to real clinical data affected by ORB and conduct a simulation study focusing on treatment effect estimation with a secondary interest in heterogeneity quantification.

stat.ME

Closed-Form Power and Sample Size Calculations for Bayes Factors

Determining an appropriate sample size is a critical element of study design, and the method used to determine it should be consistent with the planned analysis. When the planned analysis involves Bayes factor hypothesis testing, the sample size is usually desired to ensure a sufficiently high probability of obtaining a Bayes factor indicating compelling evidence for a hypothesis, given that the hypothesis is true. In practice, Bayes factor sample size determination is typically performed using computationally intensive Monte Carlo simulation. Here, we summarize alternative approaches that enable sample size determination without simulation. We show how, under approximate normality assumptions, sample sizes can be determined numerically, and provide the R package bfpwr for this purpose. Additionally, we identify conditions under which sample sizes can even be determined in closed-form, resulting in novel, easy-to-use formulas that also help foster intuition, enable asymptotic analysis, and can also be used for hybrid Bayesian/likelihoodist design. Furthermore, we show how power and sample size can be computed without simulation for more complex analysis priors, such as Jeffreys-Zellner-Siow priors or non-local normal moment priors. Case studies from medicine and psychology illustrate how researchers can use our methods to design informative yet cost-efficient studies.

stat.ME

Mixture priors for replication studies

Replication of scientific studies is important for assessing the credibility of their results. However, there is no consensus on how to quantify the extent to which a replication study replicates an original result. We propose a novel Bayesian approach for replication studies based on mixture priors. The idea is to use a mixture of the posterior distribution based on the original study and a non-informative distribution as the prior for the analysis of the replication study. The mixture weight then determines the extent to which the original and replication data are pooled. Two distinct strategies are presented: one with fixed mixture weights, and one that introduces uncertainty by assigning a prior distribution to the mixture weight itself. Furthermore, it is shown how within this framework Bayes factors can be used for formal testing of relevant scientific hypotheses, such as tests on the presence or absence of an effect or whether the mixture weight equals zero (completely discounting the original data) or one (fully pooling with the original data). To showcase the practical application of the methodology, we analyze data from three replication studies. Our findings suggest that mixture priors are a valuable and intuitive alternative to other Bayesian methods for analyzing replication studies, such as hierarchical models and power priors. We provide the free and open source R package repmix that implements the proposed methodology.

stat.ME

The assessment of replicability using the sum of p-values

Statistical significance of both the original and the replication study is a commonly used criterion to assess replication attempts, also known as the two-trials rule in drug development. However, replication studies are sometimes conducted although the original study is non-significant, in which case Type-I error rate control across both studies is no longer guaranteed. We propose an alternative method to assess replicability using the sum of p-values from the two studies. The approach provides a combined p-value and can be calibrated to control the overall Type-I error rate at the same level as the two-trials rule but allows for replication success even if the original study is non-significant. The unweighted version requires a less restrictive level of significance at replication if the original study is already convincing which facilitates sample size reductions of up to 10%. Downweighting the original study accounts for possible bias and requires a more stringent significance level and larger samples sizes at replication. Data from four large-scale replication projects are used to illustrate and compare the proposed method with the two-trials rule, meta-analysis and Fisher's combination method.

stat.AP

Outcomes truncated by death in RCTs: a simulation study on the survivor average causal effect

Continuous outcome measurements truncated by death present a challenge for the estimation of unbiased treatment effects in randomized controlled trials (RCTs). One way to deal with such situations is to estimate the survivor average causal effect (SACE), but this requires making non-testable assumptions. Motivated by an ongoing RCT in very preterm infants with intraventricular hemorrhage, we performed a simulation study to compare a SACE estimator with complete case analysis (CCA) and an analysis after multiple imputation of missing outcomes. We set up 9 scenarios combining positive, negative and no treatment effect on the outcome (cognitive development) and on survival at 2 years of age. Treatment effect estimates from all methods were compared in terms of bias, mean squared error and coverage with regard to two true treatment effects: the treatment effect on the outcome used in the simulation and the SACE, which was derived by simulation of both potential outcomes per patient. Despite targeting different estimands (principal stratum estimand, hypothetical estimand), the SACE-estimator and multiple imputation gave similar estimates of the treatment effect and efficiently reduced the bias compared to CCA. Also, both methods were relatively robust to omission of one covariate in the analysis, and thus violation of relevant assumptions. Although the SACE is not without controversy, we find it useful if mortality is inherent to the study population. Some degree of violation of the required assumptions is almost certain, but may be acceptable in practice.

stat.ME

Power priors for replication studies

The ongoing replication crisis in science has increased interest in the methodology of replication studies. We propose a novel Bayesian analysis approach using power priors: The likelihood of the original study's data is raised to the power of $α$, and then used as the prior distribution in the analysis of the replication data. Posterior distribution and Bayes factor hypothesis tests related to the power parameter $α$ quantify the degree of compatibility between the original and replication study. Inferences for other parameters, such as effect sizes, dynamically borrow information from the original study. The degree of borrowing depends on the conflict between the two studies. The practical value of the approach is illustrated on data from three replication studies, and the connection to hierarchical modeling approaches explored. We generalize the known connection between normal power priors and normal hierarchical models for fixed parameters and show that normal power prior inferences with a beta prior on the power parameter $α$ align with normal hierarchical model inferences using a generalized beta prior on the relative heterogeneity variance $I^2$. The connection illustrates that power prior modeling is unnatural from the perspective of hierarchical modeling since it corresponds to specifying priors on a relative rather than an absolute heterogeneity scale.

stat.ME

Bayesian Approaches to Designing Replication Studies

Replication studies are essential for assessing the credibility of claims from original studies. A critical aspect of designing replication studies is determining their sample size; a too small sample size may lead to inconclusive studies whereas a too large sample size may waste resources that could be allocated better in other studies. Here, we show how Bayesian approaches can be used for tackling this problem. The Bayesian framework allows researchers to combine the original data and external knowledge in a design prior distribution for the underlying parameters. Based on a design prior, predictions about the replication data can be made, and the replication sample size can be chosen to ensure a sufficiently high probability of replication success. Replication success may be defined by Bayesian or non-Bayesian criteria, and different criteria may also be combined to meet distinct stakeholders and enable conclusive inferences based on multiple analysis approaches. We investigate sample size determination in the normal-normal hierarchical model where analytical results are available and traditional sample size determination is a special case where the uncertainty on parameter values is not accounted for. We use data from a multisite replication project of social-behavioral experiments to illustrate how Bayesian approaches can help design informative and cost-effective replication studies. Our methods can be used through the R package BayesRepDesign.

stat.ME

Beyond the Two-Trials Rule

The two-trials rule for drug approval requires "at least two adequate and well-controlled studies, each convincing on its own, to establish effectiveness". This is usually implemented by requiring two significant pivotal trials and is the standard regulatory requirement to provide evidence for a new drug's efficacy. However, there is need to develop suitable alternatives to this rule for a number of reasons, including the possible availability of data from more than two trials. I consider the case of up to 3 studies and stress the importance to control the partial Type-I error rate, where only some studies have a true null effect, while maintaining the overall Type-I error rate of the two-trials rule, where all studies have a null effect. Some less-known $p$-value combination methods are useful to achieve this: Pearson's method, Edgington's method and the recently proposed harmonic mean $\chi^2$-test. I study their properties and discuss how they can be extended to a sequential assessment of success while still ensuring overall Type-I error control. I compare the different methods in terms of partial Type-I error rate, project power and the expected number of studies required. Edgington's method is eventually recommended as it is easy to implement and communicate, has only moderate partial Type-I error rate inflation but substantially increased project power.

stat.ME