SearcharxivSearch

arXiv subjects

Federico A. Bugni

Publications and source records attributed to Federico A. Bugni.

17 recordsLinked to original sources

Manipulation Testing in Boundary Discontinuity Designs

We propose the first manipulation test designed for boundary discontinuity designs (BDDs) with general boundary shapes. A BDD is a multidimensional extension of the regression discontinuity design (RDD) in which treatment assignment is determined by whether the multidimensional running variable crosses a lower-dimensional boundary set. The test avoids multivariate density estimation and builds on the observation that, in the absence of manipulation, observations near the boundary should be approximately evenly split between treatment and control within arbitrary groups defined by their projections onto the boundary. We test this implication using a collection of binomial balance tests on observations near the boundary, with groups formed by k-means clustering. We establish the asymptotic validity of the test under suitable regularity conditions. We also evaluate finite-sample performance through Monte Carlo simulations and illustrate the test in three empirical applications.

econ.EM

Testing Conditional Stochastic Dominance at Target Points

This paper introduces a test for conditional stochastic dominance between two distributions at prespecified values of a conditioning covariate, referred to as target points. The test uses a one-sided Kolmogorov--Smirnov statistic computed from induced order statistics, the outcomes attached to the conditioning observations closest to the target point, and compares it to a critical value that, given the number of neighbors, requires no resampling, kernel smoothing, or parametric assumptions. The same procedure applies whether the outcomes are continuous, discrete, or mixed, and requires only continuity of the conditional distributions in the conditioning variable. We establish asymptotic validity under two frameworks: one in which the number of neighbors is held fixed, where the induced order statistics converge to independent draws from the conditional distributions at the target point; and one in which it grows with the sample size, where we obtain an explicit rate that accommodates both an estimated target point and a data-dependent choice of the number of neighbors. We connect the test to permutation-based inference, provide a refined critical value for discrete outcomes, propose a rule for selecting the tuning parameters, and illustrate the procedure in two empirical applications whose recorded outcomes exhibit mass points. Monte Carlo simulations confirm its strong finite-sample performance.

econ.EM

On the power properties of inference for parameters with interval identified sets

This paper studies the power properties of confidence intervals (CIs) for a partially-identified parameter of interest with an interval identified set. We assume the researcher has bounds estimators needed to construct the CIs proposed by Imbens and Manski (2004), Stoye (2009), and Stoye (2020), denoted by CI_alpha^1, CI_alpha^2, CI_alpha^3, and CI_alpha^4. We also assume these bounds estimators are ``ordered'': the lower bound estimator is less than or equal to the upper bound estimator. This setup arises in economic applications involving missing data and treatment effects. Under these conditions, we establish two results. First, we show that CI_alpha^1 and CI_alpha^2 are equally powerful, and both dominate CI_alpha^3 and CI_alpha^4. Second, we consider a favorable situation in which there are two possible bounds estimators to construct these CIs, and one is more efficient than the other. One would expect that the more efficient bounds estimator yields more powerful inference. We prove that this desirable result holds for CI_alpha^1 and CI_alpha^2, but not necessarily for CI_alpha^3 or CI_alpha^4. In summary, within the class of models considered, CI_alpha^1 and CI_alpha^2 have identical power properties, and both compare favorably to CI_alpha^3 or CI_alpha^4.

econ.EM

Demand estimation without outside good shares

The BLP model is the workhorse framework for estimating demand for differentiated products using aggregate product shares. In practice, however, the share of the outside good is often unavailable. This paper studies identification and inference in the BLP model when the share of the outside good is unobserved. We show that the model is partially identified, and we derive the identified sets for the structural parameters and other quantities of economic interest. We also develop inference procedures based on moment inequalities that deliver valid confidence sets for these structural parameters and quantities of economic interest. We illustrate our results with an empirical application based on the tuna data analyzed by Gandhi et al. (2023).

econ.EM

Inference in Auctions with Many Bidders Using Transaction Prices

This paper studies inference in first-price and second-price sealed-bid auctions with many bidders, using an asymptotic framework where the number of bidders increases while the number of auctions remains fixed. Our approach enables asymptotically exact inference on key features, such as the winner's expected utility, the seller's expected revenue, and the tail of the valuation distribution, using only transaction price data. Our simulations demonstrate the accuracy of the methods in finite samples. We apply our methods to Hong Kong vehicle license auctions, focusing on high-priced, single-letter plates. Other relevant applications include online and art auctions.

econ.EM

On the Rates of Convergence of Induced Ordered Statistics and their Applications

Induced order statistics (IOS) arise when sample units are reordered according to the value of an auxiliary variable, and the associated responses are analyzed in that induced order. IOS play a central role in applications where the goal is to approximate the conditional distribution of an outcome at a fixed covariate value using observations whose covariates lie closest to that point, including regression discontinuity designs, k-nearest-neighbor methods, and distributionally robust optimization. Existing asymptotic results allow the dimension of the IOS vector to grow with the sample size only under smoothness conditions that are often too restrictive for practical data-generating processes. In particular, these conditions rule out boundary points, which are central to regression discontinuity designs. This paper develops general convergence rates for IOS under primitive and comparatively weak assumptions. We derive sharp marginal rates for the approximation of the target conditional distribution in Hellinger and total variation distances under quadratic mean differentiability and show how these marginal rates translate into joint convergence rates for the IOS vector. Our results are widely applicable: they rely on a standard smoothness condition and accommodate both interior and boundary conditioning points, as required in regression discontinuity and related settings. In the supplementary appendix, we provide complementary results under a Taylor/Holder remainder condition. Our results reveal a clear trade-off between smoothness and speed of convergence, identify regimes in which Hellinger and total variation distances behave differently, and provide explicit growth conditions on the number of nearest neighbors.

econ.EM

Decomposition and Interpretation of Treatment Effects in Settings with Delayed Outcomes

This paper studies settings where the analyst is interested in identifying and estimating the average \emph{direct} causal effect of a binary treatment on an outcome. We consider a setup in which the outcome realization does not get immediately realized after the treatment assignment, a feature that is ubiquitous in empirical settings. The period between the treatment and the realization of the outcome allows other observed actions to occur and affect the outcome. In this context, we study several regression-based estimands routinely used in empirical work to capture the average treatment effect and shed light on interpreting them in terms of ceteris paribus effects, indirect causal effects, and selection terms. We obtain three main and related takeaways under a common set of assumptions. First, the three most popular estimands do not generally satisfy what we call \emph{strong sign preservation}, in the sense that these estimands may be negative even when the treatment positively affects the outcome conditional on any possible combination of other actions. Second, the most popular regression that includes the other actions as controls satisfies strong sign preservation \emph{if and only if} these actions are mutually exclusive binary variables. Finally, we show that a linear regression that fully stratifies the other actions leads to estimands that satisfy strong sign preservation.

econ.EM

Identification and Inference on Treatment Effects under Covariate-Adaptive Randomization and Imperfect Compliance

Randomized controlled trials (RCTs) frequently utilize covariate-adaptive randomization (CAR) (e.g., stratified block randomization) and commonly suffer from imperfect compliance. This paper studies the identification and inference for the average treatment effect (ATE) and the average treatment effect on the treated (ATT) in such RCTs with a binary treatment. We first develop characterizations of the identified sets for both estimands. Since data are generally not i.i.d. under CAR, these characterizations do not follow from existing results. We then provide consistent estimators of the identified sets and asymptotically valid confidence intervals for the parameters. Our asymptotic analysis leads to concrete practical recommendations regarding how to estimate the treatment assignment probabilities that enter the estimated bounds. For the ATE bounds, using sample analog assignment frequencies is more efficient than relying on the true assignment probabilities. For the ATT bounds, the most efficient approach is to use the true assignment probability for the probabilities in the numerator and the sample analog for those in the denominator.

econ.EM

Testing homogeneity in dynamic discrete games in finite samples

The literature on dynamic discrete games often assumes that the conditional choice probabilities and the state transition probabilities are homogeneous across markets and over time. We refer to this as the "homogeneity assumption" in dynamic discrete games. This assumption enables empirical studies to estimate the game's structural parameters by pooling data from multiple markets and from many time periods. In this paper, we propose a hypothesis test to evaluate whether the homogeneity assumption holds in the data. Our hypothesis test is the result of an approximate randomization test, implemented via a Markov chain Monte Carlo (MCMC) algorithm. We show that our hypothesis test becomes valid as the (user-defined) number of MCMC draws diverges, for any fixed number of markets, time periods, and players. We apply our test to the empirical study of the U.S.\ Portland cement industry in Ryan (2012).

econ.EM

Inference under Covariate-Adaptive Randomization with Imperfect Compliance

This paper studies inference in a randomized controlled trial (RCT) with covariate-adaptive randomization (CAR) and imperfect compliance of a binary treatment. In this context, we study inference on the LATE. As in Bugni et al. (2018,2019), CAR refers to randomization schemes that first stratify according to baseline covariates and then assign treatment status so as to achieve "balance" within each stratum. In contrast to these papers, however, we allow participants of the RCT to endogenously decide to comply or not with the assigned treatment status. We study the properties of an estimator of the LATE derived from a "fully saturated" IV linear regression, i.e., a linear regression of the outcome on all indicators for all strata and their interaction with the treatment decision, with the latter instrumented with the treatment assignment. We show that the proposed LATE estimator is asymptotically normal, and we characterize its asymptotic variance in terms of primitives of the problem. We provide consistent estimators of the standard errors and asymptotically exact hypothesis tests. In the special case when the target proportion of units assigned to each treatment does not vary across strata, we can also consider two other estimators of the LATE, including the one based on the "strata fixed effects" IV linear regression, i.e., a linear regression of the outcome on indicators for all strata and the treatment decision, with the latter instrumented with the treatment assignment. Our characterization of the asymptotic variance of the LATE estimators allows us to understand the influence of the parameters of the RCT. We use this to propose strategies to minimize their asymptotic variance in a hypothetical RCT based on data from a pilot study. We illustrate the practical relevance of these results using a simulation study and an empirical application based on Dupas et al. (2018).

econ.EM

Permutation-based tests for discontinuities in event studies

We propose using a permutation test to detect discontinuities in an underlying economic model at a known cutoff point. Relative to the existing literature, we show that this test is well suited for event studies based on time-series data. The test statistic measures the distance between the empirical distribution functions of observed data in two local subsamples on the two sides of the cutoff. Critical values are computed via a standard permutation algorithm. Under a high-level condition that the observed data can be coupled by a collection of conditionally independent variables, we establish the asymptotic validity of the permutation test, allowing the sizes of the local subsamples to be either be fixed or grow to infinity. In the latter case, we also establish that the permutation test is consistent. We demonstrate that our high-level condition can be verified in a broad range of problems in the infill asymptotic time-series setting, which justifies using the permutation test to detect jumps in economic variables such as volatility, trading activity, and liquidity. These potential applications are illustrated in an empirical case study for selected FOMC announcements during the ongoing COVID-19 pandemic.

econ.EM

Permutation Tests for Equality of Distributions of Functional Data

Economic data are often generated by stochastic processes that take place in continuous time, though observations may occur only at discrete times. For example, electricity and gas consumption take place in continuous time. Data generated by a continuous time stochastic process are called functional data. This paper is concerned with comparing two or more stochastic processes that generate functional data. The data may be produced by a randomized experiment in which there are multiple treatments. The paper presents a method for testing the hypothesis that the same stochastic process generates all the functional data. The test described here applies to both functional data and multiple treatments. It is implemented as a combination of two permutation tests. This ensures that in finite samples, the true and nominal probabilities that each test rejects a correct null hypothesis are equal. The paper presents upper and lower bounds on the asymptotic power of the test under alternative hypotheses. The results of Monte Carlo experiments and an application to an experiment on billing and pricing of natural gas illustrate the usefulness of the test.

econ.EM

On the iterated estimation of dynamic discrete choice games

We study the asymptotic properties of a class of estimators of the structural parameters in dynamic discrete choice games. We consider K-stage policy iteration (PI) estimators, where K denotes the number of policy iterations employed in the estimation. This class nests several estimators proposed in the literature such as those in Aguirregabiria and Mira (2002, 2007), Pesendorfer and Schmidt-Dengler (2008), and Pakes et al. (2007). First, we establish that the K-PML estimator is consistent and asymptotically normal for all K. This complements findings in Aguirregabiria and Mira (2007), who focus on K=1 and K large enough to induce convergence of the estimator. Furthermore, we show under certain conditions that the asymptotic variance of the K-PML estimator can exhibit arbitrary patterns as a function of K. Second, we establish that the K-MD estimator is consistent and asymptotically normal for all K. For a specific weight matrix, the K-MD estimator has the same asymptotic distribution as the K-PML estimator. Our main result provides an optimal sequence of weight matrices for the K-MD estimator and shows that the optimally weighted K-MD estimator has an asymptotic distribution that is invariant to K. The invariance result is especially unexpected given the findings in Aguirregabiria and Mira (2007) for K-PML estimators. Our main result implies two new corollaries about the optimal 1-MD estimator (derived by Pesendorfer and Schmidt-Dengler (2008)). First, the optimal 1-MD estimator is optimal in the class of K-MD estimators. In other words, additional policy iterations do not provide asymptotic efficiency gains relative to the optimal 1-MD estimator. Second, the optimal 1-MD estimator is more or equally asymptotically efficient than any K-PML estimator for all K. Finally, the appendix provides appropriate conditions under which the optimal 1-MD estimator is asymptotically efficient.

econ.EM

Testing Continuity of a Density via g-order statistics in the Regression Discontinuity Design

In the regression discontinuity design (RDD), it is common practice to assess the credibility of the design by testing the continuity of the density of the running variable at the cut-off, e.g., McCrary (2008). In this paper we propose an approximate sign test for continuity of a density at a point based on the so-called g-order statistics, and study its properties under two complementary asymptotic frameworks. In the first asymptotic framework, the number q of observations local to the cut-off is fixed as the sample size n diverges to infinity, while in the second framework q diverges to infinity slowly as n diverges to infinity. Under both of these frameworks, we show that the test we propose is asymptotically valid in the sense that it has limiting rejection probability under the null hypothesis not exceeding the nominal level. More importantly, the test is easy to implement, asymptotically valid under weaker conditions than those used by competing methods, and exhibits finite sample validity under stronger conditions than those needed for its asymptotic validity. In a simulation study, we find that the approximate sign test provides good control of the rejection probability under the null hypothesis while remaining competitive under the alternative hypothesis. We finally apply our test to the design in Lee (2008), a well-known application of the RDD to study incumbency advantage.

econ.EM

Inference in partially identified models with many moment inequalities using Lasso

This paper considers inference in a partially identified moment (in)equality model with many moment inequalities. We propose a novel two-step inference procedure that combines the methods proposed by Chernozhukov, Chetverikov and Kato (2018a) (CCK18, hereafter) with a first step moment inequality selection based on the Lasso. Our method controls asymptotic size uniformly, both in underlying parameter and data distribution. Also, the power of our method compares favorably with that of the corresponding two-step method in CCK18 for large parts of the parameter space, both in theory and in simulations. Finally, we show that our Lasso-based first step can be implemented by thresholding standardized sample averages, and so it is straightforward to implement.

math.ST

Inference under Covariate-Adaptive Randomization with Multiple Treatments

This paper studies inference in randomized controlled trials with covariate-adaptive randomization when there are multiple treatments. More specifically, we study inference about the average effect of one or more treatments relative to other treatments or a control. As in Bugni et al. (2018), covariate-adaptive randomization refers to randomization schemes that first stratify according to baseline covariates and then assign treatment status so as to achieve balance within each stratum. In contrast to Bugni et al. (2018), we not only allow for multiple treatments, but further allow for the proportion of units being assigned to each of the treatments to vary across strata. We first study the properties of estimators derived from a fully saturated linear regression, i.e., a linear regression of the outcome on all interactions between indicators for each of the treatments and indicators for each of the strata. We show that tests based on these estimators using the usual heteroskedasticity-consistent estimator of the asymptotic variance are invalid; on the other hand, tests based on these estimators and suitable estimators of the asymptotic variance that we provide are exact. For the special case in which the target proportion of units being assigned to each of the treatments does not vary across strata, we additionally consider tests based on estimators derived from a linear regression with strata fixed effects, i.e., a linear regression of the outcome on indicators for each of the treatments and indicators for each of the strata. We show that tests based on these estimators using the usual heteroskedasticity-consistent estimator of the asymptotic variance are conservative, but tests based on these estimators and suitable estimators of the asymptotic variance that we provide are exact. A simulation study illustrates the practical relevance of our theoretical results.

econ.EM

Inference in Dynamic Discrete Choice Problems under Local Misspecification

Single-agent dynamic discrete choice models are typically estimated using heavily parametrized econometric frameworks, making them susceptible to model misspecification. This paper investigates how misspecification affects the results of inference in these models. Specifically, we consider a local misspecification framework in which specification errors are assumed to vanish at an arbitrary and unknown rate with the sample size. Relative to global misspecification, the local misspecification analysis has two important advantages. First, it yields tractable and general results. Second, it allows us to focus on parameters with structural interpretation, instead of "pseudo-true" parameters. We consider a general class of two-step estimators based on the K-stage sequential policy function iteration algorithm, where K denotes the number of iterations employed in the estimation. This class includes Hotz and Miller (1993)'s conditional choice probability estimator, Aguirregabiria and Mira (2002)'s pseudo-likelihood estimator, and Pesendorfer and Schmidt-Dengler (2008)'s asymptotic least squares estimator. We show that local misspecification can affect the asymptotic distribution and even the rate of convergence of these estimators. In principle, one might expect that the effect of the local misspecification could change with the number of iterations K. One of our main findings is that this is not the case, i.e., the effect of local misspecification is invariant to K. In practice, this means that researchers cannot eliminate or even alleviate problems of model misspecification by changing K.

stat.ME