SearcharxivSearch

arXiv subjects

Jake Bowers

Publications and source records attributed to Jake Bowers.

8 recordsLinked to original sources

Fully specified Bayes factors for hypothesis testing and sensitivity analysis in process tracing

In process tracing, researchers ask how strongly their evidence favors their explanation, the working theory, over a rival. Fairfield and Charman (2022) compare the two theories with a Bayes factor whose probabilities are set by hand, which Zaks (2021) argues lets researchers overstate findings. We derive those probabilities from a fully specified model of the evidence a researcher could have examined, given which observations support each theory. The working theory says most of the evidence supports it. The rival says no more than half does. The model gives two Bayes factors, neither designed to favor the working theory. The second grants the rival more benefit of the doubt and is provably the smallest value the rival's claim allows. Researchers can report how much observation bias, re-coding, or a smoking-gun weight would change a conclusion. In six published studies we reanalyze, both Bayes factors exceed 20 for two studies and not for the other four.

stat.ME

Sequential Sensitivity Analysis for Multiple Assumptions: A Framework for Understanding Racial Disparity in Police Use of Force

Inferring racial discrimination in police use of force -- the average causal effect of civilian race on use of force -- requires two assumptions about policing prior to potential use of force: that officers do not discriminate in whom they would stop (no discrimination in stops) and that, conditional on patrol context, the probability that an encounter is with a minority rather than a white civilian does not vary across encounters (no bias in encounters). As Knox et al. (2020) show, violations of the first can mask racial disparity in force. Whether it reflects discrimination in force also depends on the second. Existing sensitivity analyses address one assumption at a time. We develop a framework that varies both sequentially and apply it to NYPD Stop, Question, and Frisk data (2003--2013). Under plausible levels of discrimination in stops, we find substantial racial disparity in force. However, the conclusion that this disparity reflects discrimination is fragile to modest departures from no bias in encounters that census-based calibration suggests are demographically feasible. By jointly addressing both confounding channels, the framework reveals how they interact in ways that separate analyses cannot, contributing to understanding what generates racial disparities and how they might be addressed.

stat.ME

Randomization Tests for Distributions of Individual Treatment Effects via Combined Rank Statistics

What proportion of treated units actually benefited from an experimental intervention? What is the median or the largest individual treatment effect? This paper develops methods for answering such questions about the distribution of individual causal effects in randomized experiments. Existing approaches require the analyst to select a rank-based test statistic before observing the data. A poor choice can substantially reduce power, while searching over multiple test statistics and adjusting for multiplicity using Bonferroni correction also incurs power loss. We propose inference procedures that adaptively combine multiple rank-based statistics while maintaining finite-sample validity. For stratified experiments, we further develop weighting schemes that effectively aggregate evidence across strata of heterogeneous sizes. The resulting combined test achieves power comparable to, or exceeding, that of the best individual test, without requiring prior knowledge of the optimal statistic. When applied to a randomized experiment evaluating a teacher training program, the combined test suggests that roughly half of treated teachers benefited, whereas a single rank-based test may indicate only a small minority. Thus, the choice of test determined whether the program appears broadly successful or narrowly effective.

stat.ME

Detecting Where Effects Occur by Testing Hypotheses in Order

Experimental evaluations of public policies often randomize a new intervention within many sites or blocks. After an overall statistically significant result is reported, the natural question from a policy maker is: \emph{where} did effects occur? Standard adjustments for multiple testing answer this question with little power because they ignore how the experiment is organized: blocks nest within cohorts, sites, and districts. We organize the hypotheses in the shape of a tree that follows this administrative structure and test them top-down, stopping at any branch where the null is not rejected. A stopping rule and valid tests at each node suffice for weak control of the family-wise error rate (FWER). Whether the unadjusted procedure also controls the FWER in the strong sense depends on an \emph{error load} computable from design quantities before any data are tested; when the load exceeds one, an adaptive $\alpha$-schedule, which we prove controls the FWER on regular and irregular trees without pruning, restores control. [Correction, August 2026: an anonymous referee identified errors in the previous version. The error loads reported in the paper were computed incorrectly; corrected loads exceed 1 in all 25 block-randomized MDRC education trials at the planning effect size $d = 0.20$, so the adjustment the earlier version claimed unnecessary is in fact required. The theorem claiming FWER control under branch pruning and its switching corollary are false as stated and are withdrawn; an exact counterexample attains FWER 0.063 at $\alpha = 0.05$. The detection comparison with the Hommel procedure is under recomputation. A correction notice on page 1 details what changes and what stands; a fully corrected version is in preparation.]

stat.ME

A p-value for Process Tracing and other N=1 Studies

We introduce a method for calculating \(p\)-values to test causal hypotheses in qualitative research \emph{a la} process tracing. As in an experiment, our \(p\)-value tells us how often one would make the same or more compelling observations favoring one theory while entertaining a rival theory. We adapt Fisher's (1935) randomization-based urn model to the reality of qualitative researchers, who cannot randomize history, but can make observations about historical processes. Our test includes a method of sensitivity analysis which allows researchers to account for the possibility of observation bias, as well as a framework for representing the varying strenght of individual pieces of evidence, altoguether informing the robustness of qualitative causal inefernce. We provide simulations and replications of previously published work to illustrate how to execute our test using any type of qualitative data about events that took place within one case. This approach adds to the pluralistic turn in the use of probability theory in theory-testing process tracing by offering a simple model with provable conservatism, while relying on few assumptions the consequences of which can be directly assessed.

stat.ME

Models, Methods and Network Topology: Experimental Design for the Study of Interference

How should a network experiment be designed to achieve high statistical power? Ex- perimental treatments on networks may spread. Randomizing assignment of treatment to nodes enhances learning about the counterfactual causal effects of a social network experiment and also requires new methodology (ex. Aronow and Samii 2017a; Bow- ers et al. 2013; Toulis and Kao 2013). In this paper we show that the way in which a treatment propagates across a social network affects the statistical power of an ex- perimental design. As such, prior information regarding treatment propagation should be incorporated into the experimental design. Our findings justify reconsideration of standard practice in circumstances where units are presumed to be independent even in simple experiments: information about treatment effects is not maximized when we assign half the units to treatment and half to control. We also present an exam- ple in which statistical power depends on the extent to which the network degree of nodes is correlated with treatment assignment probability. We recommend that re- searchers think carefully about the underlying treatment propagation model motivat- ing their study in designing an experiment on a network.

stat.ME

Reasoning about Interference Between Units

If an experimental treatment is experienced by both treated and control group units, tests of hypotheses about causal effects may be difficult to conceptualize let alone execute. In this paper, we show how counterfactual causal models may be written and tested when theories suggest spillover or other network-based interference among experimental units. We show that the "no interference" assumption need not constrain scholars who have interesting questions about interference. We offer researchers the ability to model theories about how treatment given to some units may come to influence outcomes for other units. We further show how to test hypotheses about these causal effects, and we provide tools to enable researchers to assess the operating characteristics of their tests given their own models, designs, test statistics, and data. The conceptual and methodological framework we develop here is particularly applicable to social networks, but may be usefully deployed whenever a researcher wonders about interference between units. Interference between units need not be an untestable assumption; instead, interference is an opportunity to ask meaningful questions about theoretically interesting phenomena.

stat.ME

Covariate Balance in Simple, Stratified and Clustered Comparative Studies

In randomized experiments, treatment and control groups should be roughly the same--balanced--in their distributions of pretreatment variables. But how nearly so? Can descriptive comparisons meaningfully be paired with significance tests? If so, should there be several such tests, one for each pretreatment variable, or should there be a single, omnibus test? Could such a test be engineered to give easily computed $p$-values that are reliable in samples of moderate size, or would simulation be needed for reliable calibration? What new concerns are introduced by random assignment of clusters? Which tests of balance would be optimal? To address these questions, Fisher's randomization inference is applied to the question of balance. Its application suggests the reversal of published conclusions about two studies, one clinical and the other a field experiment in political participation.

stat.ME