Searcharxiv⌕ Search

arXiv subjects

Jörg Stoye

Publications and source records attributed to Jörg Stoye.

At least 19 recordsLinked to original sources

A Better Test of Choice Overload

Choice overload - in which larger choice sets are detrimental to a chooser's well-being - is potentially of great importance in the design of economic policy. Yet current evidence on its prevalence is inconclusive. We argue that existing tests are likely to be underpowered and hence that choice overload may occur more often than the literature suggests. We propose more powerful tests based on richer data and characterization theorems based on specific models, including random utility. These new approaches come with significant econometric challenges, which we show how to address. We apply our tests to new experimental data and find strong evidence of choice overload that would likely be missed using current approaches.

econ.GN↗

Local Asymptotics for Treatment Choice with Partial Identification

We provide a new asymptotic framework to derive approximately optimal treatment assignments when sampling noise from data is compounded by fundamental uncertainty due to partial identification. We recenter the reduced-form parameter around its \emph{least-favorable} configuration and consider drifting parameter sequences that yield both diminishing levels of sampling uncertainty and of partial identification. We characterize the limiting decision problem as a normal location shift model with a suitable limiting identified set. We apply our results to treatment choice problems with contaminated outcomes, to robust welfare analyses with partially identified consumer surplus, and to the problem of aggregating experimental estimates for policy adoption.

econ.EM↗

Statistical Decisions and Partial Identification: With Application to Boundary Discontinuity Design

We are delighted to respond to the excellent surveys by Cattaneo et al. (2026) and Hirano (2026). Our discussion will attempt two things: first, we show how statistical decision theory can be applied to situations with partial identification; second, we connect the surveys' themes by applying these insights to an imagined policy experiment in one of Cattaneo et al.'s (2025) applications. To do so, we lay out a stylized scenario of statistical decision making under partial identification and, drawing on our own and others' earlier work, provide a complete solution for that scenario. We then apply these results to a hypothetical reduction (modelled on actual policies) in eligibility for educational subsidies. We will see that something of interest can be said, but also that bringing the theory to the application involves some leaps of faith and leaves some questions open. This leads to the final section, where we discuss what we see as the main open challenges in statistical decision theory under partial identification.

econ.EM↗

Epsilon-Minimax Solutions of Statistical Decision Problems

A decision rule is epsilon-minimax if it is minimax up to an additive factor epsilon. We present an algorithm for provably obtaining epsilon-minimax solutions for a class of statistical decision problems. In particular, we are interested in problems where the statistician chooses randomly among I decision rules. The minimax solution of these problems admits a convex programming representation over the (I-1)-simplex. Our suggested algorithm is a well-known mirror subgradient descent routine, designed to approximately solve the convex optimization problem that defines the minimax decision rule. This iterative routine is known in the computer science literature as the hedge algorithm and is used in algorithmic game theory as a practical tool to find approximate solutions of two-person zero-sum games. We apply the suggested algorithm to different minimax problems in the econometrics literature. An empirical application to the problem of optimally selecting sites to maximize the external validity of an experimental policy evaluation illustrates the usefulness of the suggested procedure.

econ.EM↗

Robust Bayes Treatment Choice with Partial Identification

We study a class of binary treatment choice problems with partial identification through the lens of robust (multiple prior) Bayesian analysis. We use a convenient set of prior distributions to derive ex-ante and ex-post robust Bayes decision rules, both for decision makers who can randomize and for decision makers who cannot. Our main messages are as follows: First, ex-ante and ex-post robust Bayes decision rules do not agree in general, whether or not randomized rules are allowed. Second, randomized treatment assignment for some data realizations can be optimal in both ex-ante and, perhaps more surprisingly, ex-post problems. Therefore, it is usually with loss of generality to exclude randomized rules from consideration, even when regret is evaluated ex post. We apply our results to a stylized problem where a policy maker uses experimental data to choose whether to implement a new policy in a population of interest, but is concerned about the external validity of the experiment at hand (Stoye, 2012); and to the aggregation of data generated by multiple randomized control trials in different sites to make a policy choice in a population for which no experimental data are available (Manski, 2020; Ishihara and Kitagawa, 2021).

econ.EM↗

Approximate Least-Favorable Distributions and Nearly Optimal Tests via Stochastic Mirror Descent

We consider a class of hypothesis testing problems where the null hypothesis postulates $M$ distributions for the observed data, and there is only one possible distribution under the alternative. We show that one can use a stochastic mirror descent routine for convex optimization to provably obtain - after finitely many iterations - both an approximate least-favorable distribution and a nearly optimal test, in a sense we make precise. Our theoretical results yield concrete recommendations about the algorithm's implementation, including its initial condition, its step size, and the number of iterations. Importantly, our suggested algorithm can be viewed as a slight variation of the algorithm suggested by Elliott, Müller, and Watson (2015), whose theoretical performance guarantees are unknown.

econ.EM↗

Externally Valid Selection of Experimental Sites via the k-Median Problem

We present a decision-theoretic justification for viewing the question of how to best choose where to experiment in order to optimize external validity as a $k$-median problem, a popular problem in computer science and operations research. We present conditions under which minimizing the worst-case, welfare-based regret among all nonrandom schemes that select $k$ sites to experiment is approximately equal - and sometimes exactly equal - to finding the k most central vectors of baseline site-level covariates. The k-median problem can be formulated as a linear integer program. Two empirical applications illustrate the theoretical and computational benefits of the suggested procedure.

econ.EM↗

Testing Sign Congruence Between Two Parameters

We test the null hypothesis that two parameters $(μ_1,μ_2)$ have the same sign, assuming that (asymptotically) normal estimators $(\hatμ_1,\hatμ_2)$ are available. Examples of this problem include the analysis of heterogeneous treatment effects, causal interpretation of reduced-form estimands, meta-studies, and mediation analysis. A number of tests were recently proposed. We recommend a test that is simple and rejects more often than many of these recent proposals. Like all other tests in the literature, it is conservative if the truth is near $(0,0)$ and therefore also biased. To clarify whether these features are avoidable, we also provide a test that is unbiased and has exact size control on the boundary of the null hypothesis, but which has counterintuitive properties and hence we do not recommend. We use the test to improve p-values in Kowalski (2022) from information contained in that paper's main text and to establish statistical significance of some key estimates in Dippel et al. (2021).

econ.EM↗

Decision Theory for Treatment Choice Problems with Partial Identification

We apply classical statistical decision theory to a large class of treatment choice problems with partial identification. We show that, in a general class of problems with Gaussian likelihood, all decision rules are admissible; it is maximin-welfare optimal to ignore all data; and, for severe enough partial identification, there are infinitely many minimax-regret optimal decision rules, all of which sometimes randomize the policy recommendation. We uniquely characterize the minimax-regret optimal rule that least frequently randomizes, and show that, in some cases, it can outperform other minimax-regret optimal rules in terms of what we term profiled regret. We analyze the implications of our results in the aggregation of experimental estimates for policy adoption, extrapolation of Local Average Treatment Effects, and policy making in the presence of omitted variable bias.

econ.EM↗

Revealed Price Preference: Theory and Empirical Analysis

To determine the welfare implications of price changes in demand data, we introduce a revealed preference relation over prices. We show that the absence of cycles in this relation characterizes a consumer who trades off the utility of consumption against the disutility of expenditure. Our model can be applied whenever a consumer's demand over a strict subset of all available goods is being analyzed; it can also be extended to settings with discrete goods and nonlinear prices. To illustrate its use, we apply our model to a single-agent data set and to a data set with repeated cross-sections. We develop a novel test of linear hypotheses on partially identified parameters to estimate the proportion of the population who are revealed better off due to a price change in the latter application. This new technique can be used for nonparametric counterfactual analysis more broadly.

econ.EM↗

Constraint Qualifications in Partial Identification

The literature on stochastic programming typically restricts attention to problems that fulfill constraint qualifications. The literature on estimation and inference under partial identification frequently restricts the geometry of identified sets with diverse high-level assumptions. These superficially appear to be different approaches to closely related problems. We extensively analyze their relation. Among other things, we show that for partial identification through pure moment inequalities, numerous assumptions from the literature essentially coincide with the Mangasarian-Fromowitz constraint qualification. This clarifies the relation between well-known contributions, including within econometrics, and elucidates stringency, as well as ease of verification, of some high-level assumptions in seminal papers.

econ.EM↗

Bounding Infection Prevalence by Bounding Selectivity and Accuracy of Tests: With Application to Early COVID-19

I propose novel partial identification bounds on infection prevalence from information on test rate and test yield. The approach utilizes user-specified bounds on (i) test accuracy and (ii) the extent to which tests are targeted, formalized as restriction on the effect of true infection status on the odds ratio of getting tested and thereby embeddable in logit specifications. The motivating application is to the COVID-19 pandemic but the strategy may also be useful elsewhere. Evaluated on data from the pandemic's early stage, even the weakest of the novel bounds are reasonably informative. Notably, and in contrast to speculations that were widely reported at the time, they place the infection fatality rate for Italy well above the one of influenza by mid-April.

econ.EM↗

A Simple, Short, but Never-Empty Confidence Interval for Partially Identified Parameters

This paper revisits the simple, but empirically salient, problem of inference on a real-valued parameter that is partially identified through upper and lower bounds with asymptotically normal estimators. A simple confidence interval is proposed and is shown to have the following properties: - It is never empty or awkwardly short, including when the sample analog of the identified set is empty. - It is valid for a well-defined pseudotrue parameter whether or not the model is well-specified. - It involves no tuning parameters and minimal computation. Computing the interval requires concentrating out one scalar nuisance parameter. In most cases, the practical result will be simple: To achieve 95% coverage, report the union of a simple 90% (!) confidence interval for the identified set and a standard 95% confidence interval for the pseudotrue parameter. For uncorrelated estimators -- notably if bounds are estimated from distinct subsamples -- and conventional coverage levels, validity of this simple procedure can be shown analytically. The case obtains in the motivating empirical application (de Quidt, Haushofer, and Roth, 2018), in which improvement over existing inference methods is demonstrated. More generally, simulations suggest that the novel confidence interval has excellent length and size control. This is partly because, in anticipation of never being empty, the interval can be made shorter than conventional ones in relevant regions of sample space.

econ.EM↗

A Critical Assessment of Some Recent Work on COVID-19

I tentatively re-analyze data from two well-publicized studies on COVID-19, namely the Charité "viral load in children" and the Bonn "seroprevalence in Heinsberg/Gangelt" study, from information available in the preprints. The studies have the following in common: - They received worldwide attention and arguably had policy impact. - The thrusts of their findings align with the respective lead authors' (different) public stances on appropriate response to COVID-19. - Tentatively, my reading of the Gangelt study neutralizes its thrust, and my reading of the Charité study reverses it. The exercise may aid in placing these studies in the literature. With all caveats that apply to n=2 quickfire analyses based off preprints, one also wonders whether it illustrates inadvertent effects of "researcher degrees of freedom."

stat.AP↗

Confidence Intervals for Projections of Partially Identified Parameters

We propose a bootstrap-based calibrated projection procedure to build confidence intervals for single components and for smooth functions of a partially identified parameter vector in moment (in)equality models. The method controls asymptotic coverage uniformly over a large class of data generating processes. The extreme points of the calibrated projection confidence interval are obtained by extremizing the value of the function of interest subject to a proper relaxation of studentized sample analogs of the moment (in)equality conditions. The degree of relaxation, or critical level, is calibrated so that the function of theta, not theta itself, is uniformly asymptotically covered with prespecified probability. This calibration is based on repeatedly checking feasibility of linear programming problems, rendering it computationally attractive. Nonetheless, the program defining an extreme point of the confidence interval is generally nonlinear and potentially intricate. We provide an algorithm, based on the response surface method for global optimization, that approximates the solution rapidly and accurately, and we establish its rate of convergence. The algorithm is of independent interest for optimization problems with simple objectives and complicated constraints. An empirical application estimating an entry game illustrates the usefulness of the method. Monte Carlo simulations confirm the accuracy of the solution algorithm, the good statistical as well as computational performance of calibrated projection (including in comparison to other methods), and the algorithm's potential to greatly accelerate computation of other confidence intervals.

math.ST↗

Nonparametric Counterfactuals in Random Utility Models

We bound features of counterfactual choices in the nonparametric random utility model of demand, i.e. if observable choices are repeated cross-sections and one allows for unrestricted, unobserved heterogeneity. In this setting, tight bounds are developed on counterfactual discrete choice probabilities and on the expectation and c.d.f. of (functionals of) counterfactual stochastic demand.

econ.EM↗

Revealed Stochastic Preference: A One-Paragraph Proof and Generalization

McFadden and Richter (1991) and later McFadden (2005) show that the Axiom of Revealed Stochastic Preference characterizes rationalizability of choice probabilities through random utility models on finite universal choice spaces. This note proves the result in one short, elementary paragraph and extends it to set valued choice. The latter requires a different axiom than is reported in McFadden (2005).

econ.TH↗

Nonparametric Analysis of Random Utility Models

This paper develops and implements a nonparametric test of Random Utility Models. The motivating application is to test the null hypothesis that a sample of cross-sectional demand distributions was generated by a population of rational consumers. We test a necessary and sufficient condition for this that does not rely on any restriction on unobserved heterogeneity or the number of goods. We also propose and implement a control function approach to account for endogenous expenditure. An econometric result of independent interest is a test for linear inequality constraints when these are represented as the vertices of a polyhedron rather than its faces. An empirical application to the U.K. Household Expenditure Survey illustrates computational feasibility of the method in demand problems with 5 goods.

math.ST↗