SearcharxivSearch

arXiv subjects

Eric-Jan Wagenmakers

Publications and source records attributed to Eric-Jan Wagenmakers.

At least 19 recordsLinked to original sources

Meta-Analysis with JASP, Part I: Classical Approaches

Meta-analyses play a crucial part in empirical science, enabling researchers to synthesize evidence across studies and draw more precise and generalizable conclusions. Despite their importance, access to advanced meta-analytic methodology is often limited to scientists and students with considerable expertise in computer programming. To lower the barrier for adoption, we have developed the Meta-Analysis module in JASP (https://jasp-stats.org/), a free and open-source software for statistical analyses. The module offers standard and advanced meta-analytic techniques through an easy-to-use graphical user interface (GUI), allowing researchers with diverse technical backgrounds to conduct state-of-the-art analyses. This manuscript presents an overview of the meta-analytic tools implemented in the module and showcases how JASP supports a meta-analytic practice that is rigorous, relevant, and reproducible. Tutorial videos accompany the examples presented in this manuscript.

stat.ME

Extracting Bayesian Evidence from Frequentist p-Values

The $p$-value and the Bayes factor are measures of evidence that are often considered to be philosophically and mathematically incompatible: The $p$-value quantifies conflict between data and $H_0$ ("surprise"), whereas the Bayes factor quantifies the relative predictive accuracy of $H_0$ versus $H_1$ ("evidence"). We revisit Jeffreys's Approximate Bayes factor (JAB) -- a simple, largely overlooked approximation dating back to the 1930s -- which connects these two paradigms for objective hypothesis testing of the existence of an effect. Under a unit-information prior the approximation requires only the $p$-value and the effective sample size $n_\text{eff}$. We clarify the core assumptions and boundary conditions for the application of JAB and show across 704 published $t$-tests and 39 comparisons of proportions that JAB approximates objective Bayes factors remarkably well. The connection between $p$-values and JAB has a practical implication: The evidence implied by a $p$-value depends strongly on $n_\text{eff}$. Conventional verbal labels for $p$-values (e.g., "strong surprise" for .001 < $p$ < .01) correspond to similarly graded Bayes factors only around $n_\text{eff} \approx 8$; for larger samples the same $p$-value implies weaker evidence. In moderately sized to large samples, $p > .10$ can amount to moderate or even strong evidence for $H_0$. JAB offers a cheap, sample-size-sensitive supplement to $p$-values, computable from routinely reported statistics, that remains valid even under optional stopping.

stat.ME

Efficient Bayes Factor Sensitivity Analysis via Posterior Density Ratios

Bayes factor sensitivity analysis examines how the evidence for one hypothesis over another depends on the prior distribution. In complex models, the standard approach refits the model at each hyper-parameter value, and the total computational cost scales linearly in the grid size. We propose a method that recovers the entire sensitivity curve from a single additional model fit. The key identity decomposes the Bayes factor at any hyper-parameter value $γ_x$ into an ``anchor'' Bayes factor at a fixed reference $γ_0$ and a Savage--Dickey density ratio in an extended model that places a hyper-prior on $γ$. Once this extended model is fit, the Bayes factor at any $γ_x$ follows from the anchor value and a ratio of two posterior density ordinates. To approximate this ratio, we employ the importance-weighted marginal density estimator (IWMDE). Because the sensitivity parameter enters the model only through the prior distribution on the model parameters, the data likelihood cancels in the IWMDE, reducing it to a simple ratio of prior density evaluations on the MCMC draws, without any additional likelihood computation. The resulting estimator is fast, remains accurate even with small MCMC samples, and substantially outperforms kernel density estimation across the full sensitivity range. The method extends naturally to simultaneous sensitivity over multiple hyper-parameters and to Bayesian model averaging. We illustrate it on a univariate Bayesian $t$-test with exact Bayes factors for validation, a bivariate informed $t$-test, and a Bayesian model-averaged meta-analysis, obtaining accurate sensitivity curves at a fraction of the brute-force cost.

stat.ME

The Principle of Redundant Reflection

The fact that redundant information does not update a rational belief implies that rational beliefs are updated using Bayes rule. In the framework of Hild (1998a), this is true under mild conditions for discrete, continuous, and arbitrary measure spaces. We prove this result and illustrate it with two examples.

stat.ME

Meta-Analysis with JASP, Part II: Bayesian Approaches

Bayesian inference is on the rise, partly because it allows researchers to quantify parameter uncertainty, evaluate evidence for competing hypotheses, incorporate model ambiguity, and seamlessly update knowledge as information accumulates. All of these advantages apply to the meta-analytic settings; however, advanced Bayesian meta-analytic methodology is often restricted to researchers with programming experience. In order to make these tools available to a wider audience, we implemented state-of-the-art Bayesian meta-analysis methods in the Meta-Analysis module of JASP, a free and open-source statistical software package (https://jasp-stats.org/). The module allows researchers to conduct Bayesian estimation, hypothesis testing, and model averaging with models such as meta-regression, multilevel meta-analysis, and publication bias adjusted meta-analysis. Results can be interpreted using forest plots, bubble plots, and estimated marginal means. This manuscript provides an overview of the Bayesian meta-analysis tools available in JASP and demonstrates how the software enables researchers of all technical backgrounds to perform advanced Bayesian meta-analysis.

stat.ME

Fair coins tend to land on the same side they started: Evidence from 350,757 flips

Many people have flipped coins but few have stopped to ponder the statistical and physical intricacies of the process. We collected $350{,}757$ coin flips to test the counterintuitive prediction from a physics model of human coin tossing developed by Diaconis, Holmes, and Montgomery (DHM; 2007). The model asserts that when people flip an ordinary coin, it tends to land on the same side it started -- DHM estimated the probability of a same-side outcome to be about 51\%. Our data lend strong support to this precise prediction: the coins landed on the same side more often than not, $\text{Pr}(\text{same side}) = 0.508$, 95\% credible interval (CI) [$0.506$, $0.509$], $\text{BF}_{\text{same-side bias}} = 2359$. Furthermore, the data revealed considerable between-people variation in the degree of this same-side bias. Our data also confirmed the generic prediction that when people flip an ordinary coin -- with the initial side-up randomly determined -- it is equally likely to land heads or tails: $\text{Pr}(\text{heads}) = 0.500$, 95\% CI [$0.498$, $0.502$], $\text{BF}_{\text{heads-tails bias}} = 0.182$. Furthermore, this lack of heads-tails bias does not appear to vary across coins. Additional analyses revealed that the within-people same-side bias decreased as more coins were flipped, an effect that is consistent with the possibility that practice makes people flip coins in a less wobbly fashion. Our data therefore provide strong evidence that when some (but not all) people flip a fair coin, it tends to land on the same side it started.

math.HO

Power priors for replication studies

The ongoing replication crisis in science has increased interest in the methodology of replication studies. We propose a novel Bayesian analysis approach using power priors: The likelihood of the original study's data is raised to the power of $α$, and then used as the prior distribution in the analysis of the replication data. Posterior distribution and Bayes factor hypothesis tests related to the power parameter $α$ quantify the degree of compatibility between the original and replication study. Inferences for other parameters, such as effect sizes, dynamically borrow information from the original study. The degree of borrowing depends on the conflict between the two studies. The practical value of the approach is illustrated on data from three replication studies, and the connection to hierarchical modeling approaches explored. We generalize the known connection between normal power priors and normal hierarchical models for fixed parameters and show that normal power prior inferences with a beta prior on the power parameter $α$ align with normal hierarchical model inferences using a generalized beta prior on the relative heterogeneity variance $I^2$. The connection illustrates that power prior modeling is unnatural from the perspective of hierarchical modeling since it corresponds to specifying priors on a relative rather than an absolute heterogeneity scale.

stat.ME

Footprint of publication selection bias on meta-analyses in medicine, environmental sciences, psychology, and economics

Publication selection bias undermines the systematic accumulation of evidence. To assess the extent of this problem, we survey over 68,000 meta-analyses containing over 700,000 effect size estimates from medicine (67,386/597,699), environmental sciences (199/12,707), psychology (605/23,563), and economics (327/91,421). Our results indicate that meta-analyses in economics are the most severely contaminated by publication selection bias, closely followed by meta-analyses in environmental sciences and psychology, whereas meta-analyses in medicine are contaminated the least. After adjusting for publication selection bias, the median probability of the presence of an effect decreased from 99.9% to 29.7% in economics, from 98.9% to 55.7% in psychology, from 99.8% to 70.7% in environmental sciences, and from 38.0% to 29.7% in medicine. The median absolute effect sizes (in terms of standardized mean differences) decreased from d = 0.20 to d = 0.07 in economics, from d = 0.37 to d = 0.26 in psychology, from d = 0.62 to d = 0.43 in environmental sciences, and from d = 0.24 to d = 0.13 in medicine.

stat.AP

J. B. S. Haldane's Rule of Succession

After Bayes, the oldest Bayesian account of enumerative induction is given by Laplace's so-called rule of succession: if all $n$ observed instances of a phenomenon to date exhibit a given character, the probability that the next instance of that phenomenon will also exhibit the character is $\frac{n+1}{n+2}$. Laplace's rule however has the apparently counterintuitive mathematical consequence that the corresponding "universal generalization" (every future observation of this type will also exhibit that character) has zero probability. In 1932, the British scientist J. B. S. Haldane proposed an alternative rule giving a universal generalization the positive probability $\frac{n+1}{n+2} \times \frac{n+3}{n+2}$. A year later Harold Jeffreys proposed essentially the same rule in the case of a finite population. A related variant rule results in a predictive probability of $\frac{n+1}{n+2} \times \frac{n+4}{n+3}$. These arguably elegant adjustments of the original Laplacean form have the advantage that they give predictions better aligned with intuition and common sense. In this paper we discuss J. B. S. Haldane's rule and its variants, placing them in their historical context, and relating them to subsequent philosophical discussions.

stat.OT

Evidential Calibration of Confidence Intervals

We present a novel and easy-to-use method for calibrating error-rate based confidence intervals to evidence-based support intervals. Support intervals are obtained from inverting Bayes factors based on a parameter estimate and its standard error. A $k$ support interval can be interpreted as "the observed data are at least $k$ times more likely under the included parameter values than under a specified alternative". Support intervals depend on the specification of prior distributions for the parameter under the alternative, and we present several types that allow different forms of external knowledge to be encoded. We also show how prior specification can to some extent be avoided by considering a class of prior distributions and then computing so-called minimum support intervals which, for a given class of priors, have a one-to-one mapping with confidence intervals. We also illustrate how the sample size of a future study can be determined based on the concept of support. Finally, we show how the bound for the type I error rate of Bayes factors leads to a bound for the coverage of support intervals. An application to data from a clinical trial illustrates how support intervals can lead to inferences that are both intuitive and informative.

stat.ME

Normalized power priors always discount historical data

Power priors are used for incorporating historical data in Bayesian analyses by taking the likelihood of the historical data raised to the power $α$ as the prior distribution for the model parameters. The power parameter $α$ is typically unknown and assigned a prior distribution, most commonly a beta distribution. Here, we give a novel theoretical result on the resulting marginal posterior distribution of $α$ in case of the the normal and binomial model. Counterintuitively, when the current data perfectly mirror the historical data and the sample sizes from both data sets become arbitrarily large, the marginal posterior of $α$ does not converge to a point mass at $α= 1$ but approaches a distribution that hardly differs from the prior. The result implies that a complete pooling of historical and current data is impossible if a power prior with beta prior for $α$ is used.

stat.ME

Contextual aggregation and rapid updating of trial outcomes within a user-friendly open-source environment

The delayed and incomplete availability of historical findings and the lack of integrative and user-friendly software hampers the reliable interpretation of new clinical data. We developed a free, open, and user-friendly clinical trial aggregation program combining a large and representative sample of existing trial data with the latest classical and Bayesian meta-analytical models, including clear output visualizations. Our software is of particular interest for (post-graduate) educational programs (e.g., medicine, epidemiology) and global health initiatives. We demonstrate the database, interface, and plot functionality with a recent randomized controlled trial on effective epileptic seizure reduction in children treated for a parasitic brain infection. The single trial data is placed into context and we show how to interpret new results against existing knowledge instantaneously. Our program is of particular interest to those working on the contextualizing of medical findings. It may facilitate the advancement of global clinical progress as efficiently and openly as possible and simulate further bridging clinical data with the latest biostatistical models.

stat.AP

Empirical prior distributions for Bayesian meta-analyses of binary and time to event outcomes

Bayesian model-averaged meta-analysis allows quantification of evidence for both treatment effectiveness $μ$ and across-study heterogeneity $τ$. We use the Cochrane Database of Systematic Reviews to develop discipline-wide empirical prior distributions for $μ$ and $τ$ for meta-analyses of binary and time-to-event clinical trial outcomes. First, we use 50% of the database to estimate parameters of different required parametric families. Second, we use the remaining 50% of the database to select the best-performing parametric families and explore essential assumptions about the presence or absence of the treatment effectiveness and across-study heterogeneity in real data. We find that most meta-analyses of binary outcomes are more consistent with the absence of the meta-analytic effect or heterogeneity while meta-analyses of time-to-event outcomes are more consistent with the presence of the meta-analytic effect or heterogeneity. Finally, we use the complete database - with close to half a million trial outcomes - to propose specific empirical prior distributions, both for the field in general and for specific medical subdisciplines. An example from acute respiratory infections demonstrates how the proposed prior distributions can be used to conduct a Bayesian model-averaged meta-analysis in the open-source software R and JASP.

stat.ME

A general approximation to nested Bayes factors with informed priors

A staple of Bayesian model comparison and hypothesis testing, Bayes factors are often used to quantify the relative predictive performance of two rival hypotheses. The computation of Bayes factors can be challenging, however, and this has contributed to the popularity of convenient approximations such as the BIC. Unfortunately, these approximations can fail in the case of informed prior distributions. Here we address this problem by outlining an approximation to informed Bayes factors for a focal parameter $θ$. The approximation is computationally simple and requires only the maximum likelihood estimate $\hatθ$ and its standard error. The approximation uses an estimated likelihood of $θ$ and assumes that the posterior distribution for $θ$ is unaffected by the choice of prior distribution for the nuisance parameters. The resulting Bayes factor for the null hypothesis $\mathcal{H}_0: θ= θ_0$ versus the alternative hypothesis $\mathcal{H}_1: θ\sim g(θ)$ is then easily obtained using the Savage--Dickey density ratio. Three real-data examples highlight the speed and closeness of the approximation compared to bridge sampling and Laplace's method. The proposed approximation facilitates Bayesian reanalyses of standard frequentist results, encourages application of Bayesian tests with informed priors, and alleviates the computational challenges that often frustrate both Bayesian sensitivity analyses and Bayes factor design analyses. The approximation is shown to suffer under small sample sizes and when the posterior distribution of the focal parameter is substantially influenced by the prior distributions on the nuisance parameters. The proposed methodology may also be used to approximate the posterior distribution for $θ$ under $\mathcal{H}_1$.

stat.ME

Default Bayes Factors for Testing the (In)equality of Several Population Variances

Testing the (in)equality of variances is an important problem in many statistical applications. We develop default Bayes factor tests to assess the (in)equality of two or more population variances, as well as a test for whether the population variances equal a specific value. The resulting test can be used to check assumptions for commonly used procedures such as the $t$-test or ANOVA, or test substantive hypotheses concerning variances directly. We show that our Bayes factor fulfills a number of desiderata. Researchers may have directed hypotheses such as $σ_{1}^{2} > σ_{2}^{2}$, they may want to extend $\mathcal{H}_{0}$ to have a null-region, or wish to combine hypotheses about equality with hypotheses about inequality, for example $σ_{1}^{2} = σ_{2}^{2} > (σ_{3}^{2}, σ_{4}^{2})$. We extend our Bayes factor test to allow for these deviations from our proposed default and illustrate it on a number of practical examples. Our procedure is implemented in the R package $bfvartest$.

stat.ME

History and Nature of the Jeffreys-Lindley Paradox

The Jeffreys-Lindley paradox exposes a rift between Bayesian and frequentist hypothesis testing that strikes at the heart of statistical inference. Contrary to what most current literature suggests, the paradox was central to the Bayesian testing methodology developed by Sir Harold Jeffreys in the late 1930s. Jeffreys showed that the evidence against a point-null hypothesis $\mathcal{H}_0$ scales with $\sqrt{n}$ and repeatedly argued that it would therefore be mistaken to set a threshold for rejecting $\mathcal{H}_0$ at a constant multiple of the standard error. Here we summarize Jeffreys's early work on the paradox and clarify his reasons for including the $\sqrt{n}$ term. The prior distribution is seen to play a crucial role; by implicitly correcting for selection, small parameter values are identified as relatively surprising under $\mathcal{H}_1$. We highlight the general nature of the paradox by presenting both a fully frequentist and a fully Bayesian version. We also demonstrate that the paradox does not depend on assigning prior mass to a point hypothesis, as is commonly believed.

stat.ME

When Evidence and Significance Collide

Null hypothesis statistical significance testing (NHST) is the dominant approach for evaluating results from randomized controlled trials. Whereas NHST comes with long-run error rate guarantees, its main inferential tool -- the $p$-value -- is only an indirect measure of evidence against the null hypothesis. The main reason is that the $p$-value is based on the assumption the null hypothesis is true, whereas the likelihood of the data under any alternative hypothesis is ignored. If the goal is to quantify how much evidence the data provide for or against the null hypothesis it is unavoidable that an alternative hypothesis be specified (Goodman & Royall, 1988). Paradoxes arise when researchers interpret $p$-values as evidence. For instance, results that are surprising under the null may be equally surprising under a plausible alternative hypothesis, such that a $p=.045$ result (`reject the null') does not make the null any less plausible than it was before. Hence, $p$-values have been argued to overestimate the evidence against the null hypothesis. Conversely, it can be the case that statistically non-significant results (i.e., $p>.05)$ nevertheless provide some evidence in favor of the alternative hypothesis. It is therefore crucial for researchers to know when statistical significance and evidence collide, and this requires that a direct measure of evidence is computed and presented alongside the traditional $p$-value.

stat.ME

Bayes Factors for Peri-Null Hypotheses

A perennial objection against Bayes factor point-null hypothesis tests is that the point-null hypothesis is known to be false from the outset. We examine the consequences of approximating the sharp point-null hypothesis by a hazy `peri-null' hypothesis instantiated as a narrow prior distribution centered on the point of interest. The peri-null Bayes factor then equals the point-null Bayes factor multiplied by a correction term which is itself a Bayes factor. For moderate sample sizes, the correction term is relatively inconsequential; however, for large sample sizes the correction term becomes influential and causes the peri-null Bayes factor to be inconsistent and approach a limit that depends on the ratio of prior ordinates evaluated at the maximum likelihood estimate. We characterize the asymptotic behavior of the peri-null Bayes factor and briefly discuss suggestions on how to construct peri-null Bayes factor hypothesis tests that are also consistent.

math.ST