SearcharxivSearch

arXiv subjects

Xinbing Kong

Publications and source records attributed to Xinbing Kong.

11 recordsLinked to original sources

Expected Shortfall Factor Models: Common Tail Losses and Expected Returns

We develop an expected shortfall factor model (ESFM) to estimate and price common variation in the severity of lower-tail losses in large panels of asset returns. Mean factor models describe common variation in average returns, while quantile factor models describe common movements in tail thresholds. ESFM instead captures common variation in the average severity of losses below those thresholds. The model combines observed risk exposures with latent common factors. We estimate ESFM using an orthogonalized two-step procedure under which first-stage quantile estimation error has no first-order effect on the ES coefficient estimates. We establish nonasymptotic error bounds for the ES coefficients, a finite-sample Gaussian approximation, and consistent selection of the number of latent factors. Applied to a large panel of equities, ESFM uncovers common factors that react sharply to market stress and contain information not captured by mean and quantile factors. Average returns increase across portfolios sorted on ESFM exposure; high-minus-low portfolios earn annualized returns of 8.0%--11.7% and Fama--French five-factor alphas of 10.3%--15.0%. These spreads remain positive and statistically significant after conditioning separately and jointly on mean- and quantile-factor exposures. Tail-by-tail spanning tests show that ESFM factors retain significant alphas after controlling for standard traded factors and the corresponding mean and quantile factors. Adding ESFM to these benchmark factor sets increases the maximum attainable Sharpe ratio. These findings identify common loss severity as a distinct and priced dimension of downside risk.

econ.EM

One-step group factor analysis via penalized least squares

In this article, we revisit the problem of group factor analysis and propose a one-step penalized least squares method to estimate the factor loadings and factors in large-dimensional group factor models, offering a distinct alternative to the conventional two-step principal component approach. Our procedure originates from the equivalence between the group factor structure and the carefully tailored identification conditions. Leveraging this insight, we develop a tricky Lagrange multiplier formulation for a penalized least square loss function. This one-step optimization framework, combined with the fine-tuned penalty parameters, facilitates the direct derivation of central limit theorems for factor loadings, factor scores, common and local components, as well as the convergence rates for them. Our theory demonstrates that our one-step approach achieves the same rate as the two-step aggregated principal-component method and even the same limiting standard error in the balanced group panel case, but smaller limiting standard error than the two-step canonical correlation procedure. Extensive simulation studies justify the theory. Applications to U.S. house prices and CSI300 weekly returns confirm that our method identifies global factors and heterogeneous local patterns.

stat.ME

Expected Shortfall Panel Regression

Expected Shortfall (ES) is a coherent measure of tail risk that captures the average loss beyond a quantile threshold. Despite the growing literature on ES regression conditional on covariates, no existing work considers ES modeling in panel data settings where both cross-sectional and temporal dependencies are present. This paper introduces the panel ES regression model with a latent factor structure to capture cross-sectional dependence. We develop a two-stage estimation procedure robust to heavy-tailed errors, recovering the conditional quantile in the first stage and iteratively estimating the ES factor model in the second stage. Theoretically, we establish the consistency and asymptotic normality of the proposed two-step ES estimators and derive non-asymptotic error bounds for both the panel quantile and ES estimators. We also provide a non-asymptotic normal approximation for the standardized ES regression estimator, bridging asymptotic theory and finite-sample practice. Simulation evidence shows that the proposed method delivers substantial gains in both parameter estimation and factor recovery, particularly in the presence of latent tail dependence. An empirical application further indicates that the extracted ES factors carry distinct pricing information that is not captured by conventional mean or quantile-based approaches.

stat.ME

Staleness Factors and Volatility Estimation at High Frequencies

In this paper, we propose a price staleness factor model that accounts for pervasive market friction across assets and incorporates relevant covariates. Using large-panel high-frequency data, we derive the maximum likelihood estimators of the regression coefficients, the nonstationary factors, and their loading parameters. These estimators recover the time-varying price staleness probabilities. We develop asymptotic theory in which both the dimension $d$ and the sampling frequency $n$ tend to infinity. Using a local principal component analysis (LPCA) approach, we find that the efficient price co-volatilities (systematic and idiosyncratic) are biased downward due to the presence of staleness. We provide bias-corrected estimators for both the spot and integrated systematic and idiosyncratic co-volatilities, and prove that these estimators are robust to data staleness. Interestingly, besides their dependence on the dimensionality $d$, the integrated plug-in estimates converge at a rate of $n^{-1/2}$ without requiring correcting term, whereas the local PCA estimates converge at a slower rate of $n^{-1/4}$. This validates the aggregation efficiency achieved through nonlinear, nonstationary factor analysis via maximum likelihood estimation. Numerical experiments justify our theoretical findings. Empirically, we demonstrate that the staleness factor provides unique explanatory power for cross-sectional risk premia, and that the staleness correction reduces out-of-sample portfolio risk.

math.ST

Tucker Diffusion Model for High-dimensional Tensor Generation

Statistical inference on large-dimensional tensor data has been extensively studied in the literature and widely used in economics, biology, machine learning, and other fields, but how to generate a structured tensor with a target distribution is still a new problem. As profound AI generators, diffusion models have achieved remarkable success in learning complex distributions. However, their extension to generating multi-linear tensor-valued observations remains underexplored. In this work, we propose a novel Tucker diffusion model for learning high-dimensional tensor distributions. We show that the score function admits a structured decomposition under the low Tucker rank assumption, allowing it to be both accurately approximated and efficiently estimated using a carefully tailored tensor-shaped architecture named Tucker-Unet. Furthermore, the distribution of generated tensors, induced by the estimated score function, converges to the true data distribution at a rate depending on the maximum of tensor mode dimensions, thereby offering a clear theoretical advantage over the naive vectorized approach, which has a product dependence. Empirically, compared to existing approaches, the Tucker diffusion model demonstrates strong practical potential in synthetic and real-world tensor generation tasks, achieving comparable and sometimes even superior statistical performance with significantly reduced training and sampling costs.

stat.ME

Data Synchronization at High Frequencies

Asynchronous trading in high-frequency financial markets introduces significant biases into econometric analysis, distorting risk estimates and leading to suboptimal portfolio decisions. Existing synchronization methods, such as the previous-tick approach, suffer from information loss and create artificial price staleness. We introduce a novel framework that recasts the data synchronization challenge as a constrained matrix completion problem. Our approach recovers the potential matrix of high-frequency price increments by minimizing its nuclear norm -- capturing the underlying low-rank factor structure -- subject to a large-scale linear system derived from observed, asynchronous price changes. Theoretically, we prove the existence and uniqueness of our estimator and establish its convergence rate. A key theoretical insight is that our method accurately and robustly leverages information from both frequently and infrequently traded assets, overcoming a critical difficulty of efficiency loss in traditional methods. Empirically, using extensive simulations and a large panel of S&P 500 stocks, we demonstrate that our method substantially outperforms established benchmarks. It not only achieves significantly lower synchronization errors, but also corrects the bias in systematic risk estimates (i.e., eigenvalues) and the estimate of betas caused by stale prices. Crucially, portfolios constructed using our synchronized data yield consistently and economically significant higher out-of-sample Sharpe ratios. Our framework provides a powerful tool for uncovering the true dynamics of asset prices, with direct implications for high-frequency risk management, algorithmic trading, and econometric inference.

econ.EM

High-Dimensional Binary Variates: Maximum Likelihood Estimation with Nonstationary Covariates and Factors

This paper introduces a high-dimensional binary variate model that accommodates nonstationary covariates and factors, and studies their asymptotic theory. This framework encompasses scenarios where single indices are nonstationary or cointegrated. For nonstationary single indices, the maximum likelihood estimator (MLE) of the coefficients has dual convergence rates and is collectively consistent under the condition $T^{1/2}/N\to0$, as both the cross-sectional dimension $N$ and the time horizon $T$ approach infinity. The MLE of all nonstationary factors is consistent when $T^δ/N\to0$, where $δ$ depends on the link function. The limiting distributions of the factors depend on time $t$, governed by the convergence of the Hessian matrix to zero. In the case of cointegrated single indices, the MLEs of both factors and coefficients converge at a higher rate of $\min(\sqrt{N},\sqrt{T})$. A distinct feature compared to nonstationary single indices is that the dual rate of convergence of the coefficients increases from $(T^{1/4},T^{3/4})$ to $(T^{1/2},T)$. Moreover, the limiting distributions of the factors do not depend on $t$ in the cointegrated case. Monte Carlo simulations verify the accuracy of the estimates. In an empirical application, we analyze jump arrivals in financial markets using this model, extract jump arrival factors, and demonstrate their efficacy in large-cross-section asset pricing.

math.ST

Generalized Matrix Factor Model

This article introduces a nonlinear generalized matrix factor model (GMFM) that allows for mixed-type variables, extending the scope of linear matrix factor models (LMFM) that are so far limited to handling continuous variables. We introduce a novel augmented Lagrange multiplier method, equivalent to the constraint maximum likelihood estimation, and carefully tailored to be locally concave around the true factor and loading parameters. This statistically guarantees the local convexity of the negative Hessian matrix around the true parameters of the factors and loadings, which is nontrivial in the matrix factor modeling and leads to feasible central limit theorems of the estimated factors and loadings. We also theoretically establish the convergence rates of the estimated factor and loading matrices for the GMFM under general conditions that allow for correlations across samples, rows, and columns. Moreover, we provide a model selection criterion to determine the numbers of row and column factors consistently. To numerically compute the constraint maximum likelihood estimator, we provide two algorithms: two-stage alternating maximization and minorization maximization. Extensive simulation studies demonstrate GMFM's superiority in handling discrete and mixed-type variables. An empirical data analysis of the company's operating performance shows that GMFM does clustering and reconstruction well in the presence of discontinuous entries in the data matrix.

stat.ME

Matrix Factor Analysis: From Least Squares to Iterative Projection

In this article, we study large-dimensional matrix factor models and estimate the factor loading matrices and factor score matrix by minimizing square loss function. Interestingly, the resultant estimators coincide with the Projected Estimators (PE) in Yu et al.(2022), which was proposed from the perspective of simultaneous reduction of the dimensionality and the magnitudes of the idiosyncratic error matrix. In other word, we provide a least-square interpretation of the PE for matrix factor model, which parallels to the least-square interpretation of the PCA for the vector factor model. We derive the convergence rates of the theoretical minimizers under sub-Gaussian tails. Considering the robustness to the heavy tails of the idiosyncratic errors, we extend the least squares to minimizing the Huber loss function, which leads to a weighted iterative projection approach to compute and learn the parameters. We also derive the convergence rates of the theoretical minimizers of the Huber loss function under bounded $(2+ε)$th moment of the idiosyncratic errors. We conduct extensive numerical studies to investigate the empirical performance of the proposed Huber estimators relative to the state-of-the-art ones. The Huber estimators perform robustly and much better than existing ones when the data are heavy-tailed, and as a result can be used as a safe replacement in practice. An application to a Fama-French financial portfolio dataset demonstrates the empirical advantage of the Huber estimator.

stat.ME

Manifold Principle Component Analysis for Large-Dimensional Matrix Elliptical Factor Model

Matrix factor model has been growing popular in scientific fields such as econometrics, which serves as a two-way dimension reduction tool for matrix sequences. In this article, we for the first time propose the matrix elliptical factor model, which can better depict the possible heavy-tailed property of matrix-valued data especially in finance. Manifold Principle Component Analysis (MPCA) is for the first time introduced to estimate the row/column loading spaces. MPCA first performs Singular Value Decomposition (SVD)for each "local" matrix observation and then averages the local estimated spaces across all observations, while the existing ones such as 2-dimensional PCA first integrates data across observations and then does eigenvalue decomposition of the sample covariance matrices. We propose two versions of MPCA algorithms to estimate the factor loading matrices robustly, without any moment constraints on the factors and the idiosyncratic errors. Theoretical convergence rates of the corresponding estimators of the factor loading matrices, factor score matrices and common components matrices are derived under mild conditions. We also propose robust estimators of the row/column factor numbers based on the eigenvalue-ratio idea, which are proven to be consistent. Numerical studies and real example on financial returns data check the flexibility of our model and the validity of our MPCA methods.

stat.ME

Large-dimensional Factor Analysis without Moment Constraints

Large-dimensional factor model has drawn much attention in the big-data era, in order to reduce the dimensionality and extract underlying features using a few latent common factors. Conventional methods for estimating the factor model typically requires finite fourth moment of the data, which ignores the effect of heavy-tailedness and thus may result in unrobust or even inconsistent estimation of the factor space and common components. In this paper, we propose to recover the factor space by performing principal component analysis to the spatial Kendall's tau matrix instead of the sample covariance matrix. In a second step, we estimate the factor scores by the ordinary least square (OLS) regression. Theoretically, we show that under the elliptical distribution framework the factor loadings and scores as well as the common components can be estimated consistently without any moment constraint. The convergence rates of the estimated factor loadings, scores and common components are provided. The finite sample performance of the proposed procedure is assessed through thorough simulations. An analysis of a financial data set of asset returns shows the superiority of the proposed method over the classical PCA method.

stat.ME