SearcharxivSearch

arXiv subjects

Jushan Bai

Publications and source records attributed to Jushan Bai.

At least 19 recordsLinked to original sources

Fixed-$T$ Dynamic Spatial Panel Model with Common Shocks

We study a dynamic spatial panel model with observed regressors, interactive effects, and contemporaneous and lagged dependence in a large-$N$, fixed-$T$ framework. The spatial model constitutes an $N$-dimensional simultaneous-equations system. In this $N$-equation view, the interactive effects introduce $N$ unit-specific loading vectors. Estimating them individually when $T$ is fixed creates the type of incidental-parameters problem underlying Nickell bias. We use the $N$-equation view to account for spatial simultaneity through the spatial Jacobian. Crucially, however, we view the same model as a $T$-equation system with $N$ observations. Together with a spatially enriched random-loadings (SERL) specification, this $T$-equation view yields a quasi-likelihood without unit-specific incidental parameters. We propose a computationally tractable block-coordinate algorithm. Simulations show small estimation errors and generally near-nominal coverage probabilities. Applying the method to U.S. county female labor-force participation, we find substantial dynamic and spatial dependence and an important role for education in explaining the rise in female labor-force participation.

econ.EM

Causal Inference Using Factor Models

We develop a factor-model framework for causal inference in panels with policy interventions. Treatment effects are represented as structural changes in treated units' exposure to latent common shocks and, in extensions, changes in the factor process itself. The approach does not impose the standard parallel-trends restriction, accommodates one or many treated units, and targets systematic effects when unit-time idiosyncratic effects are not point identified. We provide estimation and inference under both fixed and treatment-dependent factor processes. Simulations show coverage close to nominal levels. In applications to California tobacco control and German reunification, the method produces estimates broadly consistent with synthetic control while delivering formal confidence intervals.

econ.EM

Global identification of dynamic panel models with interactive effects

We investigate the problem of global identification in dynamic panel models with interactive effects, in the large-N, fixed-T setting. While local identification, typically established via the Jacobian matrix, is well understood, global identification has remained a more elusive and challenging issue. It is commonly believed to be unachievable in this context. However, we demonstrate that the model is, in fact, globally identified for almost all configurations of the factors. Our analysis also covers models with additive fixed effects, including unit-root cases in which previous studies have reported non-identification from differenced moments. We show that, even in these settings, the level covariance structure delivers global identification.

econ.EM

Taxonomy and Estimation of Multiple Breakpoints in High-Dimensional Factor Models

This paper proposes a quasi-maximum likelihood (QML) estimator for break points in high-dimensional factor models, specifically accounting for multiple structural breaks. We begin by establishing a necessary and sufficient condition to categorize two distinct types of breaks in factor loadings: singular changes and rotational changes. The analysis of the nearly singular subsample covariance matrices of the pseudo-factors plays a key role in our approach. It allows us to demonstrate that the QML estimator precisely identifies the true breakpoint with probability tending to one for singular changes. For rotational changes, we demonstrate that the estimator exhibits stochastically bounded estimation errors, implying break fraction consistency. Furthermore, we introduce an information criterion to estimate the number of breaks, proving that it can detect the true number with probability tending to one. Monte Carlo simulations confirm the strong finite sample performance of our proposed methods. Finally, we provide an empirical example to estimate structural breakpoints in the FRED-MD dataset spanning 1959 to 2024.

econ.EM

Bayesian inference for dynamic spatial quantile models with interactive effects

With the rapid advancement of information technology and data collection systems, large-scale spatial panel data presents new methodological and computational challenges. This paper introduces a dynamic spatial panel quantile model that incorporates unobserved heterogeneity. The proposed model captures the dynamic structure of panel data, high-dimensional cross-sectional dependence, and allows for heterogeneous regression coefficients. To estimate the model, we propose a novel Bayesian Markov Chain Monte Carlo (MCMC) algorithm. Contributions to Bayesian computation include the development of quantile randomization, a new Gibbs sampler for structural parameters, and stabilization of the tail behavior of the inverse Gaussian random generator. We establish Bayesian consistency for the proposed estimation method as both the time and cross-sectional dimensions of the panel approach infinity. Monte Carlo simulations demonstrate the effectiveness of the method. Finally, we illustrate the applicability of the approach through a case study on the quantile co-movement structure of the gasoline market.

econ.EM

Efficiency of QMLE for dynamic panel data models with interactive effects

This paper studies the problem of efficient estimation of panel data models in the presence of an increasing number of incidental parameters. We formulate the dynamic panel as a simultaneous equations system, and derive the efficiency bound under the normality assumption. We then show that the Gaussian quasi-maximum likelihood estimator (QMLE) applied to the system achieves the normality efficiency bound without the normality assumption. Comparison of QMLE with the fixed effects approach is made.

econ.EM

Likelihood ratio test for structural changes in factor models

A factor model with a break in its factor loadings is observationally equivalent to a model without changes in the loadings but a change in the variance of its factors. This effectively transforms a structural change problem of high dimension into a problem of low dimension. This paper considers the likelihood ratio (LR) test for a variance change in the estimated factors. The LR test implicitly explores a special feature of the estimated factors: the pre-break and post-break variances can be a singular matrix under the alternative hypothesis, making the LR test diverging faster and thus more powerful than Wald-type tests. The better power property of the LR test is also confirmed by simulations. We also consider mean changes and multiple breaks. We apply the procedure to the factor modelling and structural change of the US employment using monthly industry-level-data.

econ.EM

Approximate Factor Models with Weaker Loadings

Pervasive cross-section dependence is increasingly recognized as a characteristic of economic data and the approximate factor model provides a useful framework for analysis. Assuming a strong factor structure where $\Lop\Lo/N^α$ is positive definite in the limit when $α=1$, early work established convergence of the principal component estimates of the factors and loadings up to a rotation matrix. This paper shows that the estimates are still consistent and asymptotically normal when $α\in(0,1]$ albeit at slower rates and under additional assumptions on the sample size. The results hold whether $α$ is constant or varies across factor loadings. The framework developed for heterogeneous loadings and the simplified proofs that can be also used in strong factor analysis are of independent interest.

econ.EM

Factor-Based Imputation of Missing Values and Covariances in Panel Data of Large Dimensions

Economists are blessed with a wealth of data for analysis, but more often than not, values in some entries of the data matrix are missing. Various methods have been proposed to handle missing observations in a few variables. We exploit the factor structure in panel data of large dimensions. Our \textsc{tall-project} algorithm first estimates the factors from a \textsc{tall} block in which data for all rows are observed, and projections of variable specific length are then used to estimate the factor loadings. A missing value is imputed as the estimated common component which we show is consistent and asymptotically normal without further iteration. Implications for using imputed data in factor augmented regressions are then discussed. To compensate for the downward bias in covariance matrices created by an omitted noise when the data point is not observed, we overlay the imputed data with re-sampled idiosyncratic residuals many times and use the average of the covariances to estimate the parameters of interest. Simulations show that the procedures have desirable finite sample properties.

econ.EM

Matrix Completion, Counterfactuals, and Factor Analysis of Missing Data

This paper proposes an imputation procedure that uses the factors estimated from a tall block along with the re-rotated loadings estimated from a wide block to impute missing values in a panel of data. Assuming that a strong factor structure holds for the full panel of data and its sub-blocks, it is shown that the common component can be consistently estimated at four different rates of convergence without requiring regularization or iteration. An asymptotic analysis of the estimation error is obtained. An application of our analysis is estimation of counterfactuals when potential outcomes have a factor structure. We study the estimation of average and individual treatment effects on the treated and establish a normal distribution theory that can be useful for hypothesis testing.

econ.EM

Quasi-maximum likelihood estimation of break point in high-dimensional factor models

This paper estimates the break point for large-dimensional factor models with a single structural break in factor loadings at a common unknown date. First, we propose a quasi-maximum likelihood (QML) estimator of the change point based on the second moments of factors, which are estimated by principal component analysis. We show that the QML estimator performs consistently when the covariance matrix of the pre- or post-break factor loading, or both, is singular. When the loading matrix undergoes a rotational type of change while the number of factors remains constant over time, the QML estimator incurs a stochastically bounded estimation error. In this case, we establish an asymptotic distribution of the QML estimator. The simulation results validate the feasibility of this estimator when used in finite samples. In addition, we demonstrate empirical applications of the proposed method by applying it to estimate the break points in a U.S. macroeconomic dataset and a stock return dataset.

econ.EM

Feasible Generalized Least Squares for Panel Data with Cross-sectional and Serial Correlations

This paper considers generalized least squares (GLS) estimation for linear panel data models. By estimating the large error covariance matrix consistently, the proposed feasible GLS (FGLS) estimator is more efficient than the ordinary least squares (OLS) in the presence of heteroskedasticity, serial, and cross-sectional correlations. To take into account the serial correlations, we employ the banding method. To take into account the cross-sectional correlations, we suggest to use the thresholding method. We establish the limiting distribution of the proposed estimator. A Monte Carlo study is considered. The proposed method is applied to an empirical application.

econ.EM

Simpler Proofs for Approximate Factor Models of Large Dimensions

Estimates of the approximate factor model are increasingly used in empirical work. Their theoretical properties, studied some twenty years ago, also laid the ground work for analysis on large dimensional panel data models with cross-section dependence. This paper presents simplified proofs for the estimates by using alternative rotation matrices, exploiting properties of low rank matrices, as well as the singular value decomposition of the data in addition to its covariance structure. These simplifications facilitate interpretation of results and provide a more friendly introduction to researchers new to the field. New results are provided to allow linear restrictions to be imposed on factor models.

econ.EM

Standard Errors for Panel Data Models with Unknown Clusters

This paper develops a new standard-error estimator for linear panel data models. The proposed estimator is robust to heteroskedasticity, serial correlation, and cross-sectional correlation of unknown forms. The serial correlation is controlled by the Newey-West method. To control for cross-sectional correlations, we propose to use the thresholding method, without assuming the clusters to be known. We establish the consistency of the proposed estimator. Monte Carlo simulations show the method works well. An empirical application is considered.

econ.EM

Robust Principal Component Analysis with Non-Sparse Errors

We show that when a high-dimensional data matrix is the sum of a low-rank matrix and a random error matrix with independent entries, the low-rank component can be consistently estimated by solving a convex minimization problem. We develop a new theoretical argument to establish consistency without assuming sparsity or the existence of any moments of the error matrix, so that fat-tailed continuous random errors such as Cauchy are allowed. The results are illustrated by simulations.

econ.EM

Principal Components and Regularized Estimation of Factor Models

It is known that the common factors in a large panel of data can be consistently estimated by the method of principal components, and principal components can be constructed by iterative least squares regressions. Replacing least squares with ridge regressions turns out to have the effect of shrinking the singular values of the common component and possibly reducing its rank. The method is used in the machine learning literature to recover low-rank matrices. We study the procedure from the perspective of estimating a minimum-rank approximate factor model. We show that the constrained factor estimates are biased but can be more efficient in terms of mean-squared errors. Rank consideration suggests a data-dependent penalty for selecting the number of factors. The new criterion is more conservative in cases when the nominal number of factors is inflated by the presence of weak factors or large measurement noise. The framework is extended to incorporate a priori linear constraints on the loadings. We provide asymptotic results that can be used to test economic hypotheses.

stat.ME

Theory and methods of panel data models with interactive effects

This paper considers the maximum likelihood estimation of panel data models with interactive effects. Motivated by applications in economics and other social sciences, a notable feature of the model is that the explanatory variables are correlated with the unobserved effects. The usual within-group estimator is inconsistent. Existing methods for consistent estimation are either designed for panel data with short time periods or are less efficient. The maximum likelihood estimator has desirable properties and is easy to implement, as illustrated by the Monte Carlo simulations. This paper develops the inferential theory for the maximum likelihood estimator, including consistency, rate of convergence and the limiting distributions. We further extend the model to include time-invariant regressors and common regressors (cross-section invariant). The regression coefficients for the time-invariant regressors are time-varying, and the coefficients for the common regressors are cross-sectionally varying.

math.ST

Statistical Inferences Using Large Estimated Covariances for Panel Data and Factor Models

While most of the convergence results in the literature on high dimensional covariance matrix are concerned about the accuracy of estimating the covariance matrix (and precision matrix), relatively less is known about the effect of estimating large covariances on statistical inferences. We study two important models: factor analysis and panel data model with interactive effects, and focus on the statistical inference and estimation efficiency of structural parameters based on large covariance estimators. For efficient estimation, both models call for a weighted principle components (WPC), which relies on a high dimensional weight matrix. This paper derives an efficient and feasible WPC using the covariance matrix estimator of Fan et al. (2013). However, we demonstrate that existing results on large covariance estimation based on absolute convergence are not suitable for statistical inferences of the structural parameters. What is needed is some weighted consistency and the associated rate of convergence, which are obtained in this paper. Finally, the proposed method is applied to the US divorce rate data. We find that the efficient WPC identifies the significant effects of divorce-law reforms on the divorce rate, and it provides more accurate estimation and tighter confidence intervals than existing methods.

math.ST