SearcharxivSearch

arXiv subjects

Hiroyuki Kasahara

Publications and source records attributed to Hiroyuki Kasahara.

17 recordsLinked to original sources

Sequential Estimation of Dynamic Discrete Choice Models with Unobserved Heterogeneity

Unobserved heterogeneity is empirically central in dynamic discrete choice models but computationally costly to incorporate: estimation requires repeatedly solving fixed-point equations for every latent type. We develop EM-NPL($q$), a unified framework combining the sequential pseudo-likelihood (NPL) and finite-mixture Expectation-Maximization (EM) algorithms, truncating the inner solver to $q$ iterations, and accommodating the Bellman, policy valuation, Euler, and Efficient Pseudo-Likelihood (EPL) equations, with step-by-step implementation guidance. For linear-in-parameters models estimated via the policy valuation or EPL equations, we establish truncation invariance: for any $q\geq 1$, EM-NPL($q$) is numerically identical to the fully converged EM-NPL estimator, so q affects computation but not statistical properties. We establish consistency, asymptotic normality, and local convergence. Truncation reduces runtime by up to 26\% for PV\_GMRES in single-agent simulations and 58\% for EPL in dynamic games. In a cola-demand application, ignoring unobserved heterogeneity understates own-price elasticities and soda-tax compensating variation by up to 85\% and 90\%.

econ.EM

Event-Study Designs for Discrete Outcomes with Latent Transition Heterogeneity

We develop an identification strategy for average treatment effects on the treated (ATT) in panel data with discrete outcomes. For such outcomes, the parallel trends assumption underlying difference-in-differences (DiD) fails in three distinct ways: mean reversion generates divergent trends when groups differ at baseline, counterfactual probabilities can leave the unit interval, and no single trend is well defined for multi-category outcomes. We replace parallel trends with \textit{transition independence}: absent treatment, transition dynamics conditional on pre-treatment outcomes would be identical between treated and control groups. To accommodate selection on persistent unobserved heterogeneity in transition dynamics, we require transition independence to hold only within latent types. Modeling outcomes as a finite mixture of Markov chains, we identify latent-type and aggregate ATTs from short panels. The framework also yields a flow decomposition of the ATT into inflow and outflow channels. In three empirical applications, our ATT estimates differ substantially from conventional DiD.

econ.EM

When Is GMM Actually LATE? Weighting Matrices and Causal Interpretation in Overidentified IV

Under heterogeneous treatment effects, the weighting matrix of overidentified IV-GMM selects the estimand, not just its precision. We characterize the selection exactly: for any parameter-free weighting-matrix map, the GMM estimand is a sum-to-one combination of instrument-specific Wald estimands, with closed-form weights and an exact non-negativity condition; efficient weighting adds a heterogeneity penalty. Continuously updated GMM exits this class through a variance-score remainder. Under positive regression dependence each Wald estimand is a convex combination of compliance-type LATEs, and under maintained validity a $J$-rejection indicates unequal Wald estimands rather than invalid instruments. We propose Representativeness Targeting (RT), which estimates a researcher-specified convex combination of the Wald estimands without imposing a common coefficient across moments; RT weights compliance types nonnegatively, attains the local asymptotic minimax bound for its target, and extends to unreachable policy targets via projection with identification-gap bounds. In Tennessee STAR, we find the $J$-test rejects the Wald-estimand equality while the heterogeneity penalty pulls the efficient-GMM estimate substantially below 2SLS; in a patent-leniency design, RT delivers a policy-relevant surrogate that standard GMM weightings miss.

econ.EM

Identification and Estimation of Production Function and Consumer Demand Function under Monopolistic Competition from Revenue Data

We establish nonparametric identification of production functions, total factor productivity (TFP), price markups, and firms' output prices and quantities, as well as consumer demand, using firm-level revenue data, without observing output quantity, in a monopolistically competitive environment with a fully nonparametric demand system. This result overturns the widely held view -- formalized by Bond, Hashemi, Kaplan, and Zoch (2021) -- that output elasticities and markups are not nonparametrically identifiable from revenue data without quantity information. Under the additional restriction that demand satisfies the homothetic single-aggregator (HSA) structure of Matsuyama and Ushchev (2017), we further nonparametrically identify the representative consumer's utility function from firm-level revenue data. This new identification result enables counterfactual welfare analysis without parametric assumptions on preferences. We propose a semiparametric estimator that is feasible for standard firm-level datasets under a Cobb--Douglas production specification. Monte Carlo simulations show that the estimator performs well, while treating revenue as output induces substantial bias. Applying the estimator to Chilean manufacturing data, we reject the CES specification in favor of HSA, and find that market power reduces welfare by approximately 3%--6% of industry revenue in the three largest manufacturing industries in 1996.

econ.EM

Gender-Specific Effects of Prenatal Famine Exposure on Educational Attainment: Accounting for Selective Mortality

Selective mortality and fertility issues are persistent challenges in estimating the fetal origin effect, with attempts to address these issues being notably scarce. Evidence further suggests that selective mortality is more pronounced in males than in females. This study investigates the causal effects of prenatal exposure to the Great Chinese Famine on educational attainment by addressing gender-specific selection bias. We compare exposed individuals with their unexposed, same-gender siblings, using a famine intensity measure based on county-year level excess death rates. Our findings reveal remarkably similar consequences for both genders: on average, famine exposure increased illiteracy rates by 4 percentage points and decreased years of schooling by 0.3 years for both males and females. These results contribute to our understanding of the long-term impacts of prenatal malnutrition, while accounting for gender-specific selection biases.

econ.GN

Semiparametric Identification of the Discount Factor and Payoff Function in Dynamic Discrete Choice Models

This paper investigates how the discount factor and payoff functions can be identified in stationary infinite-horizon dynamic discrete choice models. In single-agent models, we show that common nonparametric assumptions on per-period payoffs -- such as homogeneity of degree one, monotonicity, concavity, zero cross-differences, and complementarity -- provide identifying restrictions on the discount factor. These restrictions take the form of polynomial equalities and inequalities with degrees bounded by the cardinality of the state space. These restrictions also identify payoff functions under standard normalization at one action. In dynamic game models, we show that firm-specific discount factors can be identified using assumptions such as irrelevance of other firms' lagged actions, exchangeability, and the independence of adjustment costs from other firms' actions. Our results demonstrate that widely used nonparametric assumptions in economic analysis can provide substantial identifying power in dynamic structural models.

econ.EM

Estimating the Number of Components in Panel Data Finite Mixture Regression Models with an Application to Production Function Heterogeneity

This paper develops statistical methods for determining the number of components in panel data finite mixture regression models with regression errors independently distributed as normal or more flexible normal mixtures. We analyze the asymptotic properties of the likelihood ratio test (LRT) and information criteria (AIC and BIC) for model selection in both conditionally independent and dynamic panel settings. Unlike cross-sectional normal mixture models, we show that panel data structures eliminate higher-order degeneracy problems while retaining issues of unbounded likelihood and infinite Fisher information. Addressing these challenges, we derive the asymptotic null distribution of the LRT statistic as the maximum of random variables and develop a sequential testing procedure for consistent selection of the number of components. Our theoretical analysis also establishes the consistency of BIC and the inconsistency of AIC. Empirical application to Chilean manufacturing data reveals significant heterogeneity in production technology, with substantial variation in output elasticities of material inputs and factor-augmented technological processes within narrowly defined industries, indicating plant-specific variation in production functions beyond Hicks-neutral technological differences. These findings contrast sharply with the standard practice of assuming a homogeneous production function and highlight the necessity of accounting for unobserved plant heterogeneity in empirical production analysis.

econ.EM

Conditional Choice Probability Estimation of Dynamic Discrete Choice Models with 2-period Finite Dependence

This paper extends the work of Arcidiacono and Miller (2011, 2019) by introducing a novel characterization of finite dependence within dynamic discrete choice models, demonstrating that numerous models display 2-period finite dependence. We recast finite dependence as a problem of sequentially searching for weights and introduce a computationally efficient method for determining these weights by utilizing the Kronecker product structure embedded in state transitions. With the estimated weights, we develop a computationally attractive Conditional Choice Probability estimator with 2-period finite dependence. The computational efficacy of our proposed estimator is demonstrated through Monte Carlo simulations.

econ.EM

Testing the Number of Components in Finite Mixture Normal Regression Model with Panel Data

This paper develops the likelihood ratio-based test of the null hypothesis of a M0-component model against an alternative of (M0 + 1)-component model in the normal mixture panel regression by extending the Expectation-Maximization (EM) test of Chen and Li (2009a) and Kasahara and Shimotsu (2015) to the case of panel data. We show that, unlike the cross-sectional normal mixture, the first-order derivative of the density function for the variance parameter in the panel normal mixture is linearly independent of its second-order derivatives for the mean parameter. On the other hand, like the cross-sectional normal mixture, the likelihood ratio test statistic of the panel normal mixture is unbounded. We consider the Penalized Maximum Likelihood Estimator to deal with the unboundedness, where we obtain the data-driven penalty function via computational experiments. We derive the asymptotic distribution of the Penalized Likelihood Ratio Test (PLRT) and EM test statistics by expanding the log-likelihood function up to five times for the reparameterized parameters. The simulation experiment indicates good finite sample performance of the proposed EM test. We apply our EM test to estimate the number of production technology types for the finite mixture Cobb-Douglas production function model studied by Kasahara et al. (2022) used the panel data of the Japanese and Chilean manufacturing firms. We find the evidence of heterogeneity in elasticities of output for intermediate goods, suggesting that production function is heterogeneous across firms beyond their Hicks-neutral productivity terms.

econ.EM

Identification and Estimation of Production Function with Unobserved Heterogeneity

This paper examines the nonparametric identifiability of production functions, considering firm heterogeneity beyond Hicks-neutral technology terms. We propose a finite mixture model to account for unobserved heterogeneity in production technology and productivity growth processes. Our analysis demonstrates that the production function for each latent type can be nonparametrically identified using four periods of panel data, relying on assumptions similar to those employed in existing literature on production function and panel data identification. By analyzing Japanese plant-level panel data, we uncover significant disparities in estimated input elasticities and productivity growth processes among latent types within narrowly defined industries. We further show that neglecting unobserved heterogeneity in input elasticities may lead to substantial and systematic bias in the estimation of productivity growth.

econ.EM

A Response to Philippe Lemoine's Critique on our Paper "Causal Impact of Masks, Policies, Behavior on Early Covid-19 Pandemic in the U.S."

Recently, Phillippe Lemoine posted a critique of our paper "Causal Impact of Masks, Policies, Behavior on Early Covid-19 Pandemic in the U.S." [arXiv:2005.14168] at his post titled "Lockdowns, econometrics and the art of putting lipstick on a pig." Although Lemoine's critique appears ideologically driven and overly emotional, some of his points are worth addressing. In particular, the sensitivity of our estimation results for (i) including "masks in public spaces" and (ii) updating the data seems important critiques and, therefore, we decided to analyze the updated data ourselves. This note summarizes our findings from re-examining the updated data and responds to Phillippe Lemoine's critique on these two important points. We also briefly discuss other points Lemoine raised in his post. After analyzing the updated data, we find evidence that reinforces the conclusions reached in the original study.

stat.AP

Identification of Regression Models with a Misclassified and Endogenous Binary Regressor

We study identification in nonparametric regression models with a misclassified and endogenous binary regressor when an instrument is correlated with misclassification error. We show that the regression function is nonparametrically identified if one binary instrument variable and one binary covariate satisfy the following conditions. The instrumental variable corrects endogeneity; the instrumental variable must be correlated with the unobserved true underlying binary variable, must be uncorrelated with the error term in the outcome equation, but is allowed to be correlated with the misclassification error. The covariate corrects misclassification; this variable can be one of the regressors in the outcome equation, must be correlated with the unobserved true underlying binary variable, and must be uncorrelated with the misclassification error. We also propose a mixture-based framework for modeling unobserved heterogeneous treatment effects with a misclassified and endogenous binary regressor and show that treatment effects can be identified if the true treatment effect is related to an observed regressor and another observable variable.

econ.EM

The Association of Opening K-12 Schools with the Spread of COVID-19 in the United States: County-Level Panel Data Analysis

This paper empirically examines how the opening of K-12 schools and colleges is associated with the spread of COVID-19 using county-level panel data in the United States. Using data on foot traffic and K-12 school opening plans, we analyze how an increase in visits to schools and opening schools with different teaching methods (in-person, hybrid, and remote) is related to the 2-weeks forward growth rate of confirmed COVID-19 cases. Our debiased panel data regression analysis with a set of county dummies, interactions of state and week dummies, and other controls shows that an increase in visits to both K-12 schools and colleges is associated with a subsequent increase in case growth rates. The estimates indicate that fully opening K-12 schools with in-person learning is associated with a 5 (SE = 2) percentage points increase in the growth rate of cases. We also find that the positive association of K-12 school visits or in-person school openings with case growth is stronger for counties that do not require staff to wear masks at schools. These results have a causal interpretation in a structural model with unobserved county and time confounders. Sensitivity analysis shows that the baseline results are robust to timing assumptions and alternative specifications.

econ.GN

Nonparametric Identification of Production Function, Total Factor Productivity, and Markup from Revenue Data

Commonly used methods of production function and markup estimation assume that a firm's output quantity can be observed as data, but typical datasets contain only revenue, not output quantity. We examine the nonparametric identification of production function and markup from revenue data when a firm faces a general nonparametri demand function under imperfect competition. Under standard assumptions, we provide the constructive nonparametric identification of various firm-level objects: gross production function, total factor productivity, price markups over marginal costs, output prices, output quantities, a demand system, and a representative consumer's utility function.

econ.EM

Testing the Order of Multivariate Normal Mixture Models

Finite mixtures of multivariate normal distributions have been widely used in empirical applications in diverse fields such as statistical genetics and statistical finance. Testing the number of components in multivariate normal mixture models is a long-standing challenge even in the most important case of testing homogeneity. This paper develops likelihood-based tests of the null hypothesis of $M_0$ components against the alternative hypothesis of $M_0 + 1$ components for a general $M_0 \geq 1$. For heteroscedastic normal mixtures, we propose an EM test and derive the asymptotic distribution of the EM test statistic. For homoscedastic normal mixtures, we derive the asymptotic distribution of the likelihood ratio test statistic. We also derive the asymptotic distribution of the likelihood ratio test statistic and EM test statistic under local alternatives and show the validity of parametric bootstrap. The simulations show that the proposed test has good finite sample size and power properties.

math.ST

Asymptotic Properties of the Maximum Likelihood Estimator in Regime Switching Econometric Models

Markov regime switching models have been widely used in numerous empirical applications in economics and finance. However, the asymptotic distribution of the maximum likelihood estimator (MLE) has not been proven for some empirically popular Markov regime switching models. In particular, the asymptotic distribution of the MLE has been unknown for models in which some elements of the transition probability matrix have the value of zero, as is commonly assumed in empirical applications with models with more than two regimes. This also includes models in which the regime-specific density depends on both the current and the lagged regimes such as the seminal model of Hamilton (1989) and switching ARCH model of Hamilton and Susmel (1994). This paper shows the asymptotic normality of the MLE and consistency of the asymptotic covariance matrix estimate of these models.

math.ST

Testing the Number of Regimes in Markov Regime Switching Models

Markov regime switching models have been used in numerous empirical studies in economics and finance. However, the asymptotic distribution of the likelihood ratio test statistic for testing the number of regimes in Markov regime switching models has been an unresolved problem. This paper derives the asymptotic distribution of the likelihood ratio test statistic for testing the null hypothesis of $M_0$ regimes against the alternative hypothesis of $M_0 + 1$ regimes for any $M_0 \geq 1$ both under the null hypothesis and under local alternatives. We show that the contiguous alternatives converge to the null hypothesis at a rate of $n^{-1/8}$ in regime switching models with normal density. The asymptotic validity of the parametric bootstrap is also established.

econ.EM