SearcharxivSearch

arXiv subjects

Erkan O. Buzbas

Publications and source records attributed to Erkan O. Buzbas.

4 recordsLinked to original sources

The Difference Between "Replicable" and "Not replicable" is not Itself Scientifically Replicable

Replication studies estimate the replicability rate of scientific results by aggregating binary verdicts of experiments. Exact replications are rarely attainable, so most replication sequences are non-exact. Experiments differ in ways that matter and do not share a single data-generating process. We formalize two statistical interpretations of non-exactness. In a shared latent rate (benchmark) model, experiments are exchangeable and depend on a common random replicability rate. In a conditionally independent rates (operational) model, each experiment has its own replicability rate drawn from a population distribution. Under the benchmark model, even small variability among replicability rates induces an irreducible variance floor on the estimated mean replicability rate that no amount of replication can eliminate. Under the operational model, the degree of non-exactness is not identifiable from standard replication data, because one binary verdict per experiment carries no information about between-experiment heterogeneity. Researchers cannot tell which precision regime they are in or whether high- and low-replicability sequences can be distinguished in principle. The usual data structure cannot support reliable demarcation between "replicable" and "not replicable" results and systematically understates uncertainty, making high- and low-replicability sequences appear discriminable when they are not. We show how common sources of heterogeneity amplify these problems and demonstrate practical consequences in a reanalysis of Many Labs 4. Aggregating replicability rates across heterogeneous literatures produces averages that conflate incommensurable regimes and lack a stable interpretation. Replicability rate is not a reliable demarcation criterion. The replication crisis, if there is one, cannot be established by the methods used to declare it.

stat.AP

Openness and Reproducibility: Insights from a Model-Centric Approach

This paper investigates the conceptual relationship between openness and reproducibility using a model-centric approach, heavily informed by probability theory and statistics. We first clarify the concepts of reliability, auditability, replicability, and reproducibility--each of which denotes a potential scientific objective. Then we advance a conceptual analysis to delineate the relationship between open scientific practices and these objectives. Using the notion of an idealized experiment, we identify which components of an experiment need to be reported and which need to be repeated to achieve the relevant objective. The model-centric framework we propose aims to contribute precision and clarity to the discussions surrounding the so-called reproducibility crisis.

stat.OT

Recent selective sweeps in North American Drosophila melanogaster show signatures of soft sweeps

Rapid adaptation has been observed in numerous organisms in response to selective pressures, such as the application of pesticides and the presence of pathogens. When rapid adaptation is driven by rare alleles from the standing genetic variation or by a high population rate of de novo adaptive mutation, positive selection should commonly generate soft rather that hard selective sweeps. In a soft sweep, multiple adaptive haplotypes sweep through the population simultaneously, in contrast to hard sweeps in which only a single adaptive haplotype rises to high frequency. Current statistical methods were not designed to detect soft sweeps, and are therefore likely to miss these possibly numerous adaptive events. Here, we develop a statistical test (H12) based on haplotype homozygosity that is capable of detecting both hard and soft sweeps with similar power. We use H12 to identify multiple genomic regions that have undergone recent and strong adaptation in a population sample of fully sequenced Drosophila melanogaster strains from the Drosophila Genetic Reference Panel (DGRP). Visual inspection of the top 50 peaks revealed that multiple haplotypes are at high frequency, consistent with signatures of soft sweep. We developed a second statistic (H2/H1) that is sensitive to signatures common to soft sweeps but not hard sweeps, in order to determine whether sweeps detected by H12 can be more easily generated by hard versus soft sweeps. Surprisingly, we find that the H12 and H2/H1 values for all top 50 peaks are more easily generated by soft sweeps than hard sweeps under several evolutionary scenarios.

q-bio.PE

AABC: approximate approximate Bayesian computation when simulating a large number of data sets is computationally infeasible

Approximate Bayesian computation (ABC) methods perform inference on model-specific parameters of mechanistically motivated parametric statistical models when evaluating likelihoods is difficult. Central to the success of ABC methods is computationally inexpensive simulation of data sets from the parametric model of interest. However, when simulating data sets from a model is so computationally expensive that the posterior distribution of parameters cannot be adequately sampled by ABC, inference is not straightforward. We present approximate approximate Bayesian computation" (AABC), a class of methods that extends simulation-based inference by ABC to models in which simulating data is expensive. In AABC, we first simulate a limited number of data sets that is computationally feasible to simulate from the parametric model. We use these data sets as fixed background information to inform a non-mechanistic statistical model that approximates the correct parametric model and enables efficient simulation of a large number of data sets by Bayesian resampling methods. We show that under mild assumptions, the posterior distribution obtained by AABC converges to the posterior distribution obtained by ABC, as the number of data sets simulated from the parametric model and the sample size of the observed data set increase simultaneously. We illustrate the performance of AABC on a population-genetic model of natural selection, as well as on a model of the admixture history of hybrid populations.

stat.CO