SearcharxivSearch

arXiv subjects

Ivan A. Canay

Publications and source records attributed to Ivan A. Canay.

6 recordsLinked to original sources

Testing Conditional Stochastic Dominance at Target Points

This paper introduces a test for conditional stochastic dominance between two distributions at prespecified values of a conditioning covariate, referred to as target points. The test uses a one-sided Kolmogorov--Smirnov statistic computed from induced order statistics, the outcomes attached to the conditioning observations closest to the target point, and compares it to a critical value that, given the number of neighbors, requires no resampling, kernel smoothing, or parametric assumptions. The same procedure applies whether the outcomes are continuous, discrete, or mixed, and requires only continuity of the conditional distributions in the conditioning variable. We establish asymptotic validity under two frameworks: one in which the number of neighbors is held fixed, where the induced order statistics converge to independent draws from the conditional distributions at the target point; and one in which it grows with the sample size, where we obtain an explicit rate that accommodates both an estimated target point and a data-dependent choice of the number of neighbors. We connect the test to permutation-based inference, provide a refined critical value for discrete outcomes, propose a rule for selecting the tuning parameters, and illustrate the procedure in two empirical applications whose recorded outcomes exhibit mass points. Monte Carlo simulations confirm its strong finite-sample performance.

econ.EM

On the Rates of Convergence of Induced Ordered Statistics and their Applications

Induced order statistics (IOS) arise when sample units are reordered according to the value of an auxiliary variable, and the associated responses are analyzed in that induced order. IOS play a central role in applications where the goal is to approximate the conditional distribution of an outcome at a fixed covariate value using observations whose covariates lie closest to that point, including regression discontinuity designs, k-nearest-neighbor methods, and distributionally robust optimization. Existing asymptotic results allow the dimension of the IOS vector to grow with the sample size only under smoothness conditions that are often too restrictive for practical data-generating processes. In particular, these conditions rule out boundary points, which are central to regression discontinuity designs. This paper develops general convergence rates for IOS under primitive and comparatively weak assumptions. We derive sharp marginal rates for the approximation of the target conditional distribution in Hellinger and total variation distances under quadratic mean differentiability and show how these marginal rates translate into joint convergence rates for the IOS vector. Our results are widely applicable: they rely on a standard smoothness condition and accommodate both interior and boundary conditioning points, as required in regression discontinuity and related settings. In the supplementary appendix, we provide complementary results under a Taylor/Holder remainder condition. Our results reveal a clear trade-off between smoothness and speed of convergence, identify regimes in which Hellinger and total variation distances behave differently, and provide explicit growth conditions on the number of nearest neighbors.

econ.EM

Decomposition and Interpretation of Treatment Effects in Settings with Delayed Outcomes

This paper studies settings where the analyst is interested in identifying and estimating the average \emph{direct} causal effect of a binary treatment on an outcome. We consider a setup in which the outcome realization does not get immediately realized after the treatment assignment, a feature that is ubiquitous in empirical settings. The period between the treatment and the realization of the outcome allows other observed actions to occur and affect the outcome. In this context, we study several regression-based estimands routinely used in empirical work to capture the average treatment effect and shed light on interpreting them in terms of ceteris paribus effects, indirect causal effects, and selection terms. We obtain three main and related takeaways under a common set of assumptions. First, the three most popular estimands do not generally satisfy what we call \emph{strong sign preservation}, in the sense that these estimands may be negative even when the treatment positively affects the outcome conditional on any possible combination of other actions. Second, the most popular regression that includes the other actions as controls satisfies strong sign preservation \emph{if and only if} these actions are mutually exclusive binary variables. Finally, we show that a linear regression that fully stratifies the other actions leads to estimands that satisfy strong sign preservation.

econ.EM

On the implementation of Approximate Randomization Tests in Linear Models with a Small Number of Clusters

This paper provides a user's guide to the general theory of approximate randomization tests developed in Canay, Romano, and Shaikh (2017) when specialized to linear regressions with clustered data. An important feature of the methodology is that it applies to settings in which the number of clusters is small -- even as small as five. We provide a step-by-step algorithmic description of how to implement the test and construct confidence intervals for the parameter of interest. In doing so, we additionally present three novel results concerning the methodology: we show that the method admits an equivalent implementation based on weighted scores; we show the test and confidence intervals are invariant to whether the test statistic is studentized or not; and we prove convexity of the confidence intervals for scalar parameters. We also articulate the main requirements underlying the test, emphasizing in particular common pitfalls that researchers may encounter. Finally, we illustrate the use of the methodology with two applications that further illuminate these points. The companion {\tt R} and {\tt Stata} packages facilitate the implementation of the methodology and the replication of the empirical exercises.

econ.EM

Testing Continuity of a Density via g-order statistics in the Regression Discontinuity Design

In the regression discontinuity design (RDD), it is common practice to assess the credibility of the design by testing the continuity of the density of the running variable at the cut-off, e.g., McCrary (2008). In this paper we propose an approximate sign test for continuity of a density at a point based on the so-called g-order statistics, and study its properties under two complementary asymptotic frameworks. In the first asymptotic framework, the number q of observations local to the cut-off is fixed as the sample size n diverges to infinity, while in the second framework q diverges to infinity slowly as n diverges to infinity. Under both of these frameworks, we show that the test we propose is asymptotically valid in the sense that it has limiting rejection probability under the null hypothesis not exceeding the nominal level. More importantly, the test is easy to implement, asymptotically valid under weaker conditions than those used by competing methods, and exhibits finite sample validity under stronger conditions than those needed for its asymptotic validity. In a simulation study, we find that the approximate sign test provides good control of the rejection probability under the null hypothesis while remaining competitive under the alternative hypothesis. We finally apply our test to the design in Lee (2008), a well-known application of the RDD to study incumbency advantage.

econ.EM

Inference under Covariate-Adaptive Randomization with Multiple Treatments

This paper studies inference in randomized controlled trials with covariate-adaptive randomization when there are multiple treatments. More specifically, we study inference about the average effect of one or more treatments relative to other treatments or a control. As in Bugni et al. (2018), covariate-adaptive randomization refers to randomization schemes that first stratify according to baseline covariates and then assign treatment status so as to achieve balance within each stratum. In contrast to Bugni et al. (2018), we not only allow for multiple treatments, but further allow for the proportion of units being assigned to each of the treatments to vary across strata. We first study the properties of estimators derived from a fully saturated linear regression, i.e., a linear regression of the outcome on all interactions between indicators for each of the treatments and indicators for each of the strata. We show that tests based on these estimators using the usual heteroskedasticity-consistent estimator of the asymptotic variance are invalid; on the other hand, tests based on these estimators and suitable estimators of the asymptotic variance that we provide are exact. For the special case in which the target proportion of units being assigned to each of the treatments does not vary across strata, we additionally consider tests based on estimators derived from a linear regression with strata fixed effects, i.e., a linear regression of the outcome on indicators for each of the treatments and indicators for each of the strata. We show that tests based on these estimators using the usual heteroskedasticity-consistent estimator of the asymptotic variance are conservative, but tests based on these estimators and suitable estimators of the asymptotic variance that we provide are exact. A simulation study illustrates the practical relevance of our theoretical results.

econ.EM