SearcharxivSearch

arXiv subjects

Shakeeb Khan

Publications and source records attributed to Shakeeb Khan.

9 recordsLinked to original sources

Identification in Linear Quantile Panel Models

This paper studies identification in linear quantile panel models with unrestricted individual heterogeneity when the number of time periods is fixed and small. We impose strict exogeneity, whereby the conditional quantile restriction holds given the individual's complete regressor history and latent individual effect, but otherwise allow the disturbances to be arbitrarily dependent over time.

econ.EM

Stationary Errors and Quantile Regression in Short Panels

This paper studies a linear panel model with an unrestricted individual effect and a time- stationary idiosyncratic disturbance. We first show that stationarity is a strong restriction in a quantile model. In a linear conditional quantile specification with quantile-dependent slopes, equality of the conditional residual distributions across periods generically forces the slope coefficient to be constant over the quantile index. Thus, a stationary-error model identifies a common location coefficient rather than a collection of quantile-specific slope effects. We then develop a fixed-T estimator of this common coefficient. For each period, we run a cross- sectional quantile regression of the outcome on the full history of regressors. Stationarity makes the quantile projection of the composite individual effect and disturbance common across the period-specific regressions. Differences between diagonal and off-diagonal blocks of the resulting projection coefficients therefore identify the common slope whenever T>=2. We combine all such restrictions by a two-step minimum-distance estimator. The estimator is root-n-consistent and asymptotically normal with fixed T, permits unrestricted dependence across periods within an individual, and does not estimate the individual effects. We provide a consistent analytic covariance estimator, a cluster bootstrap, and an overidentification test of the projection restrictions implied by stationarity. Extensive Monte Carlo experiments show adequate performance under various designs.

econ.EM

Random Set Quantile Estimation of Partially Identified Discrete Response Models

Semiparametric discrete choice models are widely applied in economics, yet a fundamental tension arises when covariates are discrete as regression coefficients that are point identified under continuous regressors may become only partially identified. We show that this is not merely an identification problem but creates serious estimation pathologies. Classical estimators, including the maximum score estimator of Manski (1975), not only have population maximizers that are outer regions of the identified set (Komarova (2013)) but also converge to a random set drawn from a finite collection of deterministic regions that partition that outer region. To resolve this failure, we introduce the Random Set Quantile (RSQ) estimator which extracts the $\tau$-quantile of the classical estimator for $\tau \in (1/2,1)$. We prove this result for a class of widely used models, which includes binary/multinomial choice and discrete outcome panel data models. This construction is consistent and locally robust across the full parameter space, including precisely those configurations where classical estimators break down. A feasible implementation based on the $m$-out-of-$n$ bootstrap inherits both properties. We apply the methodology to the 2019 UK General Election, where the discrete support of Brexit-related covariates generates the partial identification our theory analyzes.

econ.EM

Inference on High Dimensional Selective Labeling Models

A class of simultaneous equation models arise in the many domains where observed binary outcomes are themselves a consequence of the existing choices of of one of the agents in the model. These models are gaining increasing interest in the computer science and machine learning literatures where they refer the potentially endogenous sample selection as the {\em selective labels} problem. Empirical settings for such models arise in fields as diverse as criminal justice, health care, and insurance. For important recent work in this area, see for example Lakkaruju et al. (2017), Kleinberg et al. (2018), and Coston et al.(2021) where the authors focus on judicial bail decisions, and where one observes the outcome of whether a defendant filed to return for their court appearance only if the judge in the case decides to release the defendant on bail. Identifying and estimating such models can be computationally challenging for two reasons. One is the nonconcavity of the bivariate likelihood function, and the other is the large number of covariates in each equation. Despite these challenges, in this paper we propose a novel distribution free estimation procedure that is computationally friendly in many covariates settings. The new method combines the semiparametric batched gradient descent algorithm introduced in Khan et al.(2023) with a novel sorting algorithms incorporated to control for selection bias. Asymptotic properties of the new procedure are established under increasing dimension conditions in both equations, and its finite sample properties are explored through a simulation study and an application using judicial bail data.

econ.EM

Sharp and Robust Estimation of Partially Identified Discrete Response Models

Semiparametric discrete choice models are widely used in a variety of practical applications. While these models are point identified in the presence of continuous covariates, they can become partially identified when covariates are discrete. In this paper we find that classical estimators, including the maximum score estimator, (Manski (1975)), loose their attractive statistical properties without point identification. First of all, they are not sharp with the estimator converging to an outer region of the identified set, (Komarova (2013)), and in many discrete designs it weakly converges to a random set. Second, they are not robust, with their distribution limit discontinuously changing with respect to the parameters of the model. We propose a novel class of estimators based on the concept of a quantile of a random set, which we show to be both sharp and robust. We demonstrate that our approach extends from cross-sectional settings to classical static and dynamic discrete panel data models.

econ.EM

Endogeneity in Weakly Separable Models without Monotonicity

We identify and estimate treatment effects when potential outcomes are weakly separable with a binary endogenous treatment. Vytlacil and Yildiz (2007) proposed an identification strategy that exploits the mean of observed outcomes, but their approach requires a monotonicity condition. In comparison, we exploit full information in the entire outcome distribution, instead of just its mean. As a result, our method does not require monotonicity and is also applicable to general settings with multiple indices. We provide examples where our approach can identify treatment effect parameters of interest whereas existing methods would fail. These include models where potential outcomes depend on multiple unobserved disturbance terms, such as a Roy model, a multinomial choice model, as well as a model with endogenous random coefficients. We establish consistency and asymptotic normality of our estimators.

econ.EM

Estimating High Dimensional Monotone Index Models by Iterative Convex Optimization1

In this paper we propose new approaches to estimating large dimensional monotone index models. This class of models has been popular in the applied and theoretical econometrics literatures as it includes discrete choice, nonparametric transformation, and duration models. A main advantage of our approach is computational. For instance, rank estimation procedures such as those proposed in Han (1987) and Cavanagh and Sherman (1998) that optimize a nonsmooth, non convex objective function are difficult to use with more than a few regressors and so limits their use in with economic data sets. For such monotone index models with increasing dimension, we propose to use a new class of estimators based on batched gradient descent (BGD) involving nonparametric methods such as kernel estimation or sieve estimation, and study their asymptotic properties. The BGD algorithm uses an iterative procedure where the key step exploits a strictly convex objective function, resulting in computational advantages. A contribution of our approach is that our model is large dimensional and semiparametric and so does not require the use of parametric distributional assumptions.

econ.EM

Identification and Estimation of Weakly Separable Models Without Monotonicity

We study the identification and estimation of treatment effect parameters in weakly separable models. In their seminal work, Vytlacil and Yildiz (2007) showed how to identify and estimate the average treatment effect of a dummy endogenous variable when the outcome is weakly separable in a single index. Their identification result builds on a monotonicity condition with respect to this single index. In comparison, we consider similar weakly separable models with multiple indices, and relax the monotonicity condition for identification. Unlike Vytlacil and Yildiz (2007), we exploit the full information in the distribution of the outcome variable, instead of just its mean. Indeed, when the outcome distribution function is more informative than the mean, our method is applicable to more general settings than theirs; in particular we do not rely on their monotonicity assumption and at the same time we also allow for multiple indices. To illustrate the advantage of our approach, we provide examples of models where our approach can identify parameters of interest whereas existing methods would fail. These examples include models with multiple unobserved disturbance terms such as the Roy model and multinomial choice models with dummy endogenous variables, as well as potential outcome models with endogenous random coefficients. Our method is easy to implement and can be applied to a wide class of models. We establish standard asymptotic properties such as consistency and asymptotic normality.

econ.EM

Informational Content of Factor Structures in Simultaneous Binary Response Models

We study the informational content of factor structures in discrete triangular systems. Factor structures have been employed in a variety of settings in cross sectional and panel data models, and in this paper we formally quantify their identifying power in a bivariate system often employed in the treatment effects literature. Our main findings are that imposing a factor structure yields point identification of parameters of interest, such as the coefficient associated with the endogenous regressor in the outcome equation, under weaker assumptions than usually required in these models. In particular, we show that a "non-standard" exclusion restriction that requires an explanatory variable in the outcome equation to be excluded from the treatment equation is no longer necessary for identification, even in cases where all of the regressors from the outcome equation are discrete. We also establish identification of the coefficient of the endogenous regressor in models with more general factor structures, in situations where one has access to at least two continuous measurements of the common factor.

econ.EM