SearcharxivSearch

arXiv subjects

Liangjun Su

Publications and source records attributed to Liangjun Su.

12 recordsLinked to original sources

Two-Way Mean Group Estimators for Heterogeneous Panel Models with Fixed T

We consider a correlated random coefficient panel data model with two-way fixed effects and interactive fixed effects in a fixed T framework. We propose a two-way mean group (TW-MG) estimator for the expected value of the slope coefficient and propose a leave-one-out jackknife method for valid inference. We also consider a pooled estimator and provide a Hausman-type test for poolability. Simulations demonstrate the excellent performance of our estimators and inference methods in finite samples. We apply our new methods to two datasets to examine the relationship between health-care expenditure and income, and estimate a production function.

econ.EM

A Robust Residual-Based Test for Structural Changes in Factor Models

In this paper, we propose an easy-to-implement residual-based specification testing procedure for detecting structural changes in factor models, which is powerful against both smooth and abrupt structural changes with unknown break dates. The proposed test is robust against the over-specified number of factors, and serially and crosssectionally correlated error processes. A new central limit theorem is given for the quadratic forms of panel data with dependence over both dimensions, thereby filling a gap in the literature. We establish the asymptotic properties of the proposed test statistic, and accordingly develop a simulation-based scheme to select critical value in order to improve finite sample performance. Through extensive simulations and a real-world application, we confirm our theoretical results and demonstrate that the proposed test exhibits desirable size and power in practice.

econ.EM

High-dimensional inference on jumps in nonparametric time series regression models

We study simultaneous inference on jumps in the conditional mean functions of a high-dimensional collection of heterogeneous nonparametric time series, where the number of series may exceed the sample size and the data may exhibit strong cross-sectional dependence. The jump depends on one specific covariate, and we allow the regression function to vary with additional latent variables. We propose two uniform tests: one for the existence of jumps and one for their homogeneity across series. We derive a simple closed-form approximation to the covariance structure of the jump estimators and establish a high-dimensional Gaussian approximation showing that, owing to the localized construction of the statistics, the maximum of the studentized jumps is approximated by the maximum of independent Gaussians. The cross-sectional dependence is thus asymptotically negligible for critical values, even under strong (e.g., factor) dependence, and the approximation requires estimating only the variance for each series. For pronounced cross-sectional dependence, a dependence-aware refinement restores the off-diagonal covariances, improving finite-sample size and power. Simulations show accurate size and reasonable power under both cross-sectional and serial dependence, and two empirical applications reveal significant non-smooth effects.

econ.EM

Panel Data Models with Time-Varying Latent Group Structures

This paper considers a linear panel model with interactive fixed effects and unobserved individual and time heterogeneities that are captured by some latent group structures and an unknown structural break, respectively. To enhance realism the model may have different numbers of groups and/or different group memberships before and after the break. With the preliminary nuclear-norm-regularized estimation followed by row- and column-wise linear regressions, we estimate the break point based on the idea of binary segmentation and the latent group structures together with the number of groups before and after the break by sequential testing K-means algorithm simultaneously. It is shown that the break point, the number of groups and the group memberships can each be estimated correctly with probability approaching one. Asymptotic distributions of the estimators of the slope coefficients are established. Monte Carlo simulations demonstrate excellent finite sample performance for the proposed estimation algorithm. An empirical application to real house price data across 377 Metropolitan Statistical Areas in the US from 1975 to 2014 suggests the presence both of structural breaks and of changes in group membership.

econ.EM

Low-rank Panel Quantile Regression: Estimation and Inference

In this paper, we propose a class of low-rank panel quantile regression models which allow for unobserved slope heterogeneity over both individuals and time. We estimate the heterogeneous intercept and slope matrices via nuclear norm regularization followed by sample splitting, row- and column-wise quantile regressions and debiasing. We show that the estimators of the factors and factor loadings associated with the intercept and slope matrices are asymptotically normally distributed. In addition, we develop two specification tests: one for the null hypothesis that the slope coefficient is a constant over time and/or individuals under the case that true rank of slope matrix equals one, and the other for the null hypothesis that the slope coefficient exhibits an additive structure under the case that the true rank of slope matrix equals two. We illustrate the finite sample performance of estimation and inference via Monte Carlo simulations and real datasets.

econ.EM

A One-Covariate-at-a-Time Method for Nonparametric Additive Models

This paper proposes a one-covariate-at-a-time multiple testing (OCMT) approach to choose significant variables in high-dimensional nonparametric additive regression models. Similarly to Chudik, Kapetanios and Pesaran (2018), we consider the statistical significance of individual nonparametric additive components one at a time and take into account the multiple testing nature of the problem. One-stage and multiple-stage procedures are both considered. The former works well in terms of the true positive rate only if the marginal effects of all signals are strong enough; the latter helps to pick up hidden signals that have weak marginal effects. Simulations demonstrate the good finite sample performance of the proposed procedures. As an empirical application, we use the OCMT procedure on a dataset we extracted from the Longitudinal Survey on Rural Urban Migration in China. We find that our procedure works well in terms of the out-of-sample forecast root mean square errors, compared with competing methods.

econ.EM

Interactive Effects Panel Data Models with General Factors and Regressors

This paper considers a model with general regressors and unobservable factors. An estimator based on iterated principal components is proposed, which is shown to be not only asymptotically normal and oracle efficient, but under certain conditions also free of the otherwise so common asymptotic incidental parameters bias. Interestingly, the conditions required to achieve unbiasedness become weaker the stronger the trends in the factors, and if the trending is strong enough unbiasedness comes at no cost at all. In particular, the approach does not require any knowledge of how many factors there are, or whether they are deterministic or stochastic. The order of integration of the factors is also treated as unknown, as is the order of integration of the regressors, which means that there is no need to pre-test for unit roots, or to decide on which deterministic terms to include in the model.

econ.EM

L2-Relaxation: With Applications to Forecast Combination and Portfolio Analysis

This paper tackles forecast combination with many forecasts or minimum variance portfolio selection with many assets. A novel convex problem called L2-relaxation is proposed. In contrast to standard formulations, L2-relaxation minimizes the squared Euclidean norm of the weight vector subject to a set of relaxed linear inequality constraints. The magnitude of relaxation, controlled by a tuning parameter, balances the bias and variance. When the variance-covariance (VC) matrix of the individual forecast errors or financial assets exhibits latent group structures -- a block equicorrelation matrix plus a VC for idiosyncratic noises, the solution to L2-relaxation delivers roughly equal within-group weights. Optimality of the new method is established under the asymptotic framework when the number of the cross-sectional units $N$ potentially grows much faster than the time dimension $T$. Excellent finite sample performance of our method is demonstrated in Monte Carlo simulations. Its wide applicability is highlighted in three real data examples concerning empirical applications of microeconomics, macroeconomics, and finance.

econ.EM

Detecting Latent Communities in Network Formation Models

This paper proposes a logistic undirected network formation model which allows for assortative matching on observed individual characteristics and the presence of edge-wise fixed effects. We model the coefficients of observed characteristics to have a latent community structure and the edge-wise fixed effects to be of low rank. We propose a multi-step estimation procedure involving nuclear norm regularization, sample splitting, iterative logistic regression and spectral clustering to detect the latent communities. We show that the latent communities can be exactly recovered when the expected degree of the network is of order log n or higher, where n is the number of nodes in the network. The finite sample performance of the new estimation and inference methods is illustrated through both simulated and real datasets.

econ.EM

Determining the Number of Communities in Degree-corrected Stochastic Block Models

We propose to estimate the number of communities in degree-corrected stochastic block models based on a pseudo likelihood ratio statistic. To this end, we introduce a method that combines spectral clustering with binary segmentation. This approach guarantees an upper bound for the pseudo likelihood ratio statistic when the model is over-fitted. We also derive its limiting distribution when the model is under-fitted. Based on these properties, we establish the consistency of our estimator for the true number of communities. Developing these theoretical properties require a mild condition on the average degrees -- growing at a rate no slower than log(n), where n is the number of nodes. Our proposed method is further illustrated by simulation studies and analysis of real-world networks. The numerical results show that our approach has satisfactory performance when the network is semi-dense.

stat.ME

Strong Consistency of Spectral Clustering for Stochastic Block Models

In this paper we prove the strong consistency of several methods based on the spectral clustering techniques that are widely used to study the community detection problem in stochastic block models (SBMs). We show that under some weak conditions on the minimal degree, the number of communities, and the eigenvalues of the probability block matrix, the K-means algorithm applied to the eigenvectors of the graph Laplacian associated with its first few largest eigenvalues can classify all individuals into the true community uniformly correctly almost surely. Extensions to both regularized spectral clustering and degree-corrected SBMs are also considered. We illustrate the performance of different methods on simulated networks.

stat.ME

Non-separable Models with High-dimensional Data

This paper studies non-separable models with a continuous treatment when the dimension of the control variables is high and potentially larger than the effective sample size. We propose a three-step estimation procedure to estimate the average, quantile, and marginal treatment effects. In the first stage we estimate the conditional mean, distribution, and density objects by penalized local least squares, penalized local maximum likelihood estimation, and numerical differentiation, respectively, where control variables are selected via a localized method of L1-penalization at each value of the continuous treatment. In the second stage we estimate the average and marginal distribution of the potential outcome via the plug-in principle. In the third stage, we estimate the quantile and marginal treatment effects by inverting the estimated distribution function and using the local linear regression, respectively. We study the asymptotic properties of these estimators and propose a weighted-bootstrap method for inference. Using simulated and real datasets, we demonstrate that the proposed estimators perform well in finite samples.

stat.ME