Searcharxiv⌕ Search

arXiv subjects

Stefan Faridani

Publications and source records attributed to Stefan Faridani.

5 recordsLinked to original sources

Designing Spatial Treatments

Spatial treatments are interventions assigned to locations potentially distinct from those of the responding units. We study their optimal design under a general model in which a unit's response diminishes with distance to a treated site. Our estimand of interest is an ``uncontaminated'' effect equal to the average impact of a single intervention site over all hypothetical sites. We propose a novel design based on a Matérn point process which separates treatments by a distance of at least $r$. A larger choice of $r$ reduces bias by separating interventions but increases variance by reducing their numerosity. We choose $r$ to maximize the rate of convergence of a Horvitz-Thompson estimator and prove that this is minimax rate-optimal. We provide weak conditions under which the estimator is asymptotically normal and propose a variance estimator.

econ.EM↗

How Replicable Are Statistically Significant Findings?

In the empirical sciences, significance thresholds often determine whether findings are treated as evidence of an effect. This paper studies how likely findings that just meet conventional significance thresholds are to remain significant in replications of the same sample size. To answer this question, we estimate the expected replication probability conditional on a given p-value among published studies for experimental economics, psychology, and social science. We validate this measure by showing it accurately predicts actual replication outcomes, outperforming prediction markets. A finding with a p-value of 0.05 has an expected replication probability ranging from 0.10 to 0.25 across fields. Low replicability reflects low power in original studies rather than publication bias. We then develop a nonparametric estimator and apply it to economics literatures that use larger samples, finding higher but still low replication probabilities. These results indicate that statistical significance in a single study provides only suggestive evidence of an effect. Stronger conclusions require cumulative evidence.

econ.EM↗

Testing for Underpowered Literatures

How many experimental studies would have come to different conclusions had they been run on larger samples? I show how to estimate the expected number of statistically significant results that a set of experiments would have reported had their sample sizes all been counterfactually increased. The proposed deconvolution estimator is asymptotically normal and adjusts for publication bias. Unlike related methods, this approach requires no assumptions of any kind about the distribution of true intervention treatment effects and allows for point masses. Simulations find good coverage even when the t-score is only approximately normal. An application to randomized trials (RCTs) published in economics journals finds that doubling every sample would increase the power of t-tests by 7.2 percentage points on average. This effect is smaller than for non-RCTs and comparable to systematic replications in laboratory psychology where previous studies enabled more accurate power calculations. This suggests that RCTs are on average relatively insensitive to sample size increases. Research funders who wish to raise power should generally consider sponsoring better-measured and higher quality experiments -- rather than only larger ones.

econ.EM↗

When is p-hacking detectable?

We show that some forms of p-hacking cannot be detected by examining the histogram of t-statistics or their p-values. Even when p-hacking is detectable, standard tests may lack power. We propose a novel test that detects every form of selective reporting that is detectable from the distribution of reported t-statistics. Our test statistic is the distance between the smoothed empirical t-curve and the set of possible honest distributions. This projection test is sharp and can only be evaded by selective reporting that also evades all other valid tests of restrictions on the t-curve. We also show how to avoid spurious rejections caused by some benign distortions in the t-curve. Applying the test to the Brodeur et al. (2020) meta-dataset, we find that the t-curves for RCTs and IVs are more distorted than could arise by chance, (de)rounding, or the Student-t approximation.

econ.EM↗

Linear estimation of global average treatment effects

We study estimation of and inference for the average causal effect of treating every member of a population, as opposed to none, using an experiment that treats only some. Considering settings where spillovers can occur between any pair of units and decay slowly with distance, we derive the minimax rate over all linear estimators and experimental designs, which increases with the spatial rate of spillover decay. This rate of convergence can be achieved using an inverse probability weighting estimator when randomization clusters are large, but not otherwise. If the causal model is linear, however, an OLS-based estimator converges faster than IPW when clusters are small and is consistent even under unit-level randomization. We provide methods for radius selection and inference and apply these to the cash transfer experiment studied by Egger et al. (2022), obtaining a 22% larger estimated effect on consumption.

econ.EM↗