SearcharxivSearch

arXiv subjects

Nick Doudchenko

Publications and source records attributed to Nick Doudchenko.

5 recordsLinked to original sources

Integer Programming for Generalized Causal Bootstrap Designs

In experimental causal inference, we distinguish between two sources of uncertainty: design uncertainty, due to the treatment assignment mechanism, and sampling uncertainty, when the sample is drawn from a super-population. This distinction matters in settings with small fixed samples and heterogeneous treatment effects, as in geographical experiments. The standard bootstrap procedure most often used by practitioners primarily estimates sampling uncertainty, and the causal bootstrap procedure, which accounts for design uncertainty, was developed for the completely randomized design and the difference-in-means estimator, whereas non-standard designs and estimators are often used in these low-power regimes. We address this gap by proposing an integer program which computes numerically the worst-case copula used as an input to the causal bootstrap method, in a wide range of settings. Specifically, we prove the asymptotic validity of our approach for unconfounded, conditionally unconfounded, and and individualistic with bounded confoundedness assignments, as well as generalizing to any linear-in-treatment and quadratic-in-treatment estimators. We demonstrate the refined confidence intervals achieved through simulations of small geographical experiments.

stat.ME

Supergeo Design: Generalized Matching for Geographic Experiments

We propose a generalization of the standard matched pairs design in which experimental units (often geographic regions or geos) may be combined into larger units/regions called "supergeos" in order to improve the average matching quality. Unlike optimal matched pairs design which can be found in polynomial time (Lu et al. 2011), this generalized matching problem is NP-hard. We formulate it as a mixed-integer program (MIP) and show that experimental design obtained by solving this MIP can often provide a significant improvement over the standard design regardless of whether the treatment effects are homogeneous or heterogeneous. Furthermore, we present the conditions under which trimming techniques that often improve performance in the case of homogeneous effects (Chen and Au, 2022), may lead to biased estimates and show that the proposed design does not introduce such bias. We use empirical studies based on real-world advertising data to illustrate these findings.

stat.ME

Synthetic Design: An Optimization Approach to Experimental Design with Synthetic Controls

We investigate the optimal design of experimental studies that have pre-treatment outcome data available. The average treatment effect is estimated as the difference between the weighted average outcomes of the treated and control units. A number of commonly used approaches fit this formulation, including the difference-in-means estimator and a variety of synthetic-control techniques. We propose several methods for choosing the set of treated units in conjunction with the weights. Observing the NP-hardness of the problem, we introduce a mixed-integer programming formulation which selects both the treatment and control sets and unit weightings. We prove that these proposed approaches lead to qualitatively different experimental units being selected for treatment. We use simulations based on publicly available data from the US Bureau of Labor Statistics that show improvements in terms of mean squared error and statistical power when compared to simple and commonly used alternatives such as randomized trials.

stat.ME

Estimation of Discrete Choice Models: A Machine Learning Approach

In this paper we propose a new method of estimation for discrete choice demand models when individual level data are available. The method employs a two-step procedure. Step 1 predicts the choice probabilities as functions of the observed individual level characteristics. Step 2 estimates the structural parameters of the model using the estimated choice probabilities at a particular point of interest and the moment restrictions. In essence, the method uses nonparametric approximation (followed by) moment estimation. Hence the name---NAME. We use simulations to compare the performance of NAME with the standard methodology. We find that our method improves precision as well as convergence time. We supplement the analysis by providing the large sample properties of the proposed estimator.

stat.AP

Causal Inference with Bipartite Designs

Bipartite experiments are a recent object of study in causal inference, whereby treatment is applied to one set of units and outcomes of interest are measured on a different set of units. These experiments are particularly useful in settings where strong interference effects occur between units of a bipartite graph. In market experiments for example, assigning treatment at the seller-level and measuring outcomes at the buyer-level (or vice-versa) may lead to causal models that better account for the interference that naturally occurs between buyers and sellers. While bipartite experiments have been shown to improve the estimation of causal effects in certain settings, the analysis must be done carefully so as to not introduce unnecessary bias. We leverage the generalized propensity score literature to show that we can obtain unbiased estimates of causal effects for bipartite experiments under a standard set of assumptions. We also discuss the construction of confidence sets with proper coverage probabilities. We evaluate these methods using a bipartite graph from a publicly available dataset studied in previous work on bipartite experiments, showing through simulations a significant bias reduction and improved coverage.

stat.ME