SearcharxivSearch

arXiv subjects

Thomas Tao Yang

Publications and source records attributed to Thomas Tao Yang.

6 recordsLinked to original sources

High Dimensional Binary Choice Model with Unknown Heteroskedasticity or Instrumental Variables

This paper proposes a new method for estimating high-dimensional binary choice models. We consider a semiparametric model that places no distributional assumptions on the error term, allows for heteroskedastic errors, and permits endogenous regressors. Our approaches extend the special regressor estimator originally proposed by Lewbel (2000). This estimator becomes impractical in high-dimensional settings due to the curse of dimensionality associated with high-dimensional conditional density estimation. To overcome this challenge, we introduce an innovative data-driven dimension reduction method for nonparametric kernel estimators, which constitutes the main contribution of this work. The method combines distance covariance-based screening with cross-validation (CV) procedures, making special regressor estimation feasible in high dimensions. Using this new feasible conditional density estimator, we address variable and moment (instrumental variable) selection problems for these models. We apply penalized least squares (LS) and generalized method of moments (GMM) estimators with an L1 penalty. A comprehensive analysis of the oracle and asymptotic properties of these estimators is provided. Finally, through Monte Carlo simulations and an empirical study on the migration intentions of rural Chinese residents, we demonstrate the effectiveness of our proposed methods in finite sample settings.

econ.EM

Semiparametric Discrete Choice Models for Bundles

We propose two approaches to estimate semiparametric discrete choice models for bundles. Our first approach is a kernel-weighted rank estimator based on a matching-based identification strategy. We establish its complete asymptotic properties and prove the validity of the nonparametric bootstrap for inference. We then introduce a new multi-index least absolute deviations (LAD) estimator as an alternative, of which the main advantage is its capacity to estimate preference parameters on both alternative- and agent-specific regressors. Both methods can account for arbitrary correlation in disturbances across choices, with the former also allowing for interpersonal heteroskedasticity. We also demonstrate that the identification strategy underlying these procedures can be extended naturally to panel data settings, producing an analogous localized maximum score estimator and a LAD estimator for estimating bundle choice models with fixed effects. We derive the limiting distribution of the former and verify the validity of the numerical bootstrap as an inference tool. All our proposed methods can be applied to general multi-index models. Monte Carlo experiments show that they perform well in finite samples.

econ.EM

Revisiting Panel Data Discrete Choice Models with Lagged Dependent Variables

This paper revisits the identification and estimation of a class of semiparametric (distribution-free) panel data binary choice models with lagged dependent variables, exogenous covariates, and entity fixed effects. We provide a novel identification strategy, using an "identification at infinity" argument. In contrast with the celebrated Honore and Kyriazidou (2000), our method permits time trends of any form and does not suffer from the "curse of dimensionality". We propose an easily implementable conditional maximum score estimator. The asymptotic properties of the proposed estimator are fully characterized. A small-scale Monte Carlo study demonstrates that our approach performs satisfactorily in finite samples. We illustrate the usefulness of our method by presenting an empirical application to enrollment in private hospital insurance using the Household, Income and Labour Dynamics in Australia (HILDA) Survey data.

econ.EM

A One-Covariate-at-a-Time Method for Nonparametric Additive Models

This paper proposes a one-covariate-at-a-time multiple testing (OCMT) approach to choose significant variables in high-dimensional nonparametric additive regression models. Similarly to Chudik, Kapetanios and Pesaran (2018), we consider the statistical significance of individual nonparametric additive components one at a time and take into account the multiple testing nature of the problem. One-stage and multiple-stage procedures are both considered. The former works well in terms of the true positive rate only if the marginal effects of all signals are strong enough; the latter helps to pick up hidden signals that have weak marginal effects. Simulations demonstrate the good finite sample performance of the proposed procedures. As an empirical application, we use the OCMT procedure on a dataset we extracted from the Longitudinal Survey on Rural Urban Migration in China. We find that our procedure works well in terms of the out-of-sample forecast root mean square errors, compared with competing methods.

econ.EM

Semiparametric Estimation of Dynamic Binary Choice Panel Data Models

We propose a new approach to the semiparametric analysis of panel data binary choice models with fixed effects and dynamics (lagged dependent variables). The model we consider has the same random utility framework as in Honore and Kyriazidou (2000). We demonstrate that, with additional serial dependence conditions on the process of deterministic utility and tail restrictions on the error distribution, the (point) identification of the model can proceed in two steps, and only requires matching the value of an index function of explanatory variables over time, as opposed to that of each explanatory variable. Our identification approach motivates an easily implementable, two-step maximum score (2SMS) procedure -- producing estimators whose rates of convergence, in contrast to Honore and Kyriazidou's (2000) methods, are independent of the model dimension. We then derive the asymptotic properties of the 2SMS procedure and propose bootstrap-based distributional approximations for inference. Monte Carlo evidence indicates that our procedure performs adequately in finite samples.

econ.EM

Quasi-Bayesian Inference for Production Frontiers

We propose a quasi-Bayesian method to conduct inference for the production frontier. This approach combines multiple first-stage extreme quantile estimates by the quasi-Bayesian method to produce the point estimate and confidence interval for the production frontier. We show the asymptotic properties of the proposed estimator and the validity of the inference procedure. The finite sample performance of our method is illustrated through simulations and an empirical application.

stat.ME