SearcharxivSearch

arXiv subjects

Frederik Aust

Publications and source records attributed to Frederik Aust.

5 recordsLinked to original sources

Extracting Bayesian Evidence from Frequentist p-Values

The $p$-value and the Bayes factor are measures of evidence that are often considered to be philosophically and mathematically incompatible: The $p$-value quantifies conflict between data and $H_0$ ("surprise"), whereas the Bayes factor quantifies the relative predictive accuracy of $H_0$ versus $H_1$ ("evidence"). We revisit Jeffreys's Approximate Bayes factor (JAB) -- a simple, largely overlooked approximation dating back to the 1930s -- which connects these two paradigms for objective hypothesis testing of the existence of an effect. Under a unit-information prior the approximation requires only the $p$-value and the effective sample size $n_\text{eff}$. We clarify the core assumptions and boundary conditions for the application of JAB and show across 704 published $t$-tests and 39 comparisons of proportions that JAB approximates objective Bayes factors remarkably well. The connection between $p$-values and JAB has a practical implication: The evidence implied by a $p$-value depends strongly on $n_\text{eff}$. Conventional verbal labels for $p$-values (e.g., "strong surprise" for .001 < $p$ < .01) correspond to similarly graded Bayes factors only around $n_\text{eff} \approx 8$; for larger samples the same $p$-value implies weaker evidence. In moderately sized to large samples, $p > .10$ can amount to moderate or even strong evidence for $H_0$. JAB offers a cheap, sample-size-sensitive supplement to $p$-values, computable from routinely reported statistics, that remains valid even under optional stopping.

stat.ME

Fair coins tend to land on the same side they started: Evidence from 350,757 flips

Many people have flipped coins but few have stopped to ponder the statistical and physical intricacies of the process. We collected $350{,}757$ coin flips to test the counterintuitive prediction from a physics model of human coin tossing developed by Diaconis, Holmes, and Montgomery (DHM; 2007). The model asserts that when people flip an ordinary coin, it tends to land on the same side it started -- DHM estimated the probability of a same-side outcome to be about 51\%. Our data lend strong support to this precise prediction: the coins landed on the same side more often than not, $\text{Pr}(\text{same side}) = 0.508$, 95\% credible interval (CI) [$0.506$, $0.509$], $\text{BF}_{\text{same-side bias}} = 2359$. Furthermore, the data revealed considerable between-people variation in the degree of this same-side bias. Our data also confirmed the generic prediction that when people flip an ordinary coin -- with the initial side-up randomly determined -- it is equally likely to land heads or tails: $\text{Pr}(\text{heads}) = 0.500$, 95\% CI [$0.498$, $0.502$], $\text{BF}_{\text{heads-tails bias}} = 0.182$. Furthermore, this lack of heads-tails bias does not appear to vary across coins. Additional analyses revealed that the within-people same-side bias decreased as more coins were flipped, an effect that is consistent with the possibility that practice makes people flip coins in a less wobbly fashion. Our data therefore provide strong evidence that when some (but not all) people flip a fair coin, it tends to land on the same side it started.

math.HO

Power priors for replication studies

The ongoing replication crisis in science has increased interest in the methodology of replication studies. We propose a novel Bayesian analysis approach using power priors: The likelihood of the original study's data is raised to the power of $α$, and then used as the prior distribution in the analysis of the replication data. Posterior distribution and Bayes factor hypothesis tests related to the power parameter $α$ quantify the degree of compatibility between the original and replication study. Inferences for other parameters, such as effect sizes, dynamically borrow information from the original study. The degree of borrowing depends on the conflict between the two studies. The practical value of the approach is illustrated on data from three replication studies, and the connection to hierarchical modeling approaches explored. We generalize the known connection between normal power priors and normal hierarchical models for fixed parameters and show that normal power prior inferences with a beta prior on the power parameter $α$ align with normal hierarchical model inferences using a generalized beta prior on the relative heterogeneity variance $I^2$. The connection illustrates that power prior modeling is unnatural from the perspective of hierarchical modeling since it corresponds to specifying priors on a relative rather than an absolute heterogeneity scale.

stat.ME

Normalized power priors always discount historical data

Power priors are used for incorporating historical data in Bayesian analyses by taking the likelihood of the historical data raised to the power $α$ as the prior distribution for the model parameters. The power parameter $α$ is typically unknown and assigned a prior distribution, most commonly a beta distribution. Here, we give a novel theoretical result on the resulting marginal posterior distribution of $α$ in case of the the normal and binomial model. Counterintuitively, when the current data perfectly mirror the historical data and the sample sizes from both data sets become arbitrarily large, the marginal posterior of $α$ does not converge to a point mass at $α= 1$ but approaches a distribution that hardly differs from the prior. The result implies that a complete pooling of historical and current data is impossible if a power prior with beta prior for $α$ is used.

stat.ME

Informed Bayesian survival analysis

We overview Bayesian estimation, hypothesis testing, and model-averaging and illustrate how they benefit parametric survival analysis. We contrast the Bayesian framework to the currently dominant frequentist approach and highlight advantages, such as seamless incorporation of historical data, continuous monitoring of evidence, and incorporating uncertainty about the true data generating process. We illustrate the application of the Bayesian approaches on an example data set from a colon cancer trial. We compare the Bayesian parametric survival analysis and frequentist models with AIC/BIC model selection in fixed-n and sequential designs with a simulation study. In the example data set, the Bayesian framework provided evidence for the absence of a positive treatment effect on disease-free survival in patients with resected colon cancer. Furthermore, the Bayesian sequential analysis would have terminated the trial 10.3 months earlier than the standard frequentist analysis. In a simulation study with sequential designs, the Bayesian framework on average reached a decision in almost half the time required by the frequentist counterparts, while maintaining the same power, and an appropriate false-positive rate. Under model misspecification, the Bayesian framework resulted in higher false-negative rate compared to the frequentist counterparts, which resulted in a higher proportion of undecided trials. In fixed-n designs, the Bayesian framework showed slightly higher power, slightly elevated error rates, and lower bias and RMSE when estimating treatment effects in small samples. We have made the analytic approach readily available in RoBSA R package. The outlined Bayesian framework provides several benefits when applied to parametric survival analyses. It uses data more efficiently, is capable of greatly shortening the length of clinical trials, and provides a richer set of inferences.

stat.ME