SearcharxivSearch

arXiv subjects

Yanrong Yang

Publications and source records attributed to Yanrong Yang.

At least 19 recordsLinked to original sources

Distributionally Robust PCA with Data-Adaptive Wasserstein Geometry

We develop a distributionally robust formulation of principal component analysis that minimizes worst-case reconstruction risk over distributions lying within a Wasserstein neighborhood of the empirical measure. The Wasserstein neighborhood, viewed as an ambiguity set of distributions, is adaptively calibrated through a transport matrix $G$ to capture heterogeneous uncertainty across dimensions. The homogeneous case, in which G is a scalar multiple of the identity matrix, recovers classical PCA. Under a general transport matrix G, we derive a dual characterization of the associated minimax optimization problem and introduce a tractable surrogate objective function consisting of the square-root empirical reconstruction error plus a geometry-dependent residual exposure penalty. The exact and surrogate estimators are shown to be consistent for the population PCA subspace and asymptotically equivalent at the projector level. The transport geometry is allowed to be data adaptive, while the Wasserstein radius is calibrated via robust Wasserstein profile inference, yielding a data-driven radius of order $n^{-1/2}$. Comprehensive theoretical guarantees are established, including consistency and local Grassmannian asymptotics exhibiting an explicit Wasserstein-induced drift determined by the limiting transport geometry and calibration level. Numerical experiments and a real-data application demonstrate that the proposed method can substantially improve finite-sample out-of-sample performance under structured covariance shifts, moderate contamination, and certain same-distribution regimes.

math.ST

Attribution of Spurious Factors from High-Dimensional Functional Time Series

This article explores a general factor structure for high-dimensional nonstationary functional time series, encompassing a wide range of factor models studied in the existing literature. We investigate the asymptotic spectral behaviors of the sample covariance operator under this general data structure. A novel fundamental sufficient condition, formulated in terms of a newly introduced effective rank tailored to this setup, is established under which empirical eigen-analysis yields spurious results, rendering sample eigenvalues and eigenvectors unreliable for accurately recovering the underlying factor structure. This generalizes the results of Onatski and Wang [2021] from typical high-dimensional time series (HDTS) to the more intricate functional framework. The newly defined effective rank is rigorously analyzed through a decomposition of the effects attributable to functional factor loadings and functional factors. Contrary to the findings in the HDTS setting, empirical eigen-analysis of models with only a small number of strong non-stationary factors may still produce spurious limits in the functional framework. Therefore, additional caution is warranted when applying covariance-based statistical methods to potentially nonstationary functional data. Simulation studies are performed to determine conditions under which spurious limits occur. Real data analysis on age-specific mortality rate data from multiple locations is conducted for evidence of spurious factors induced by empirical eigen-analysis.

stat.ME

Sustainable Investment: ESG Impacts on Large Portfolio

This paper investigates the impact of environmental, social, and governance (ESG) constraint on a regularized mean-variance (MV) portfolio optimization problem in a large-dimensional setting, in which a positive definite regularization matrix is imposed on the sample covariance matrix. We first derive the asymptotic results for the out-of-sample (OOS) Sharpe ratio (SR) of the proposed portfolio, which help quantify the impact of imposing an ESG-level constraint as well as the effect of estimation error arising from the sample mean estimation of the assets' ESG score. Furthermore, to study the influence of the choices of the regularization matrix, we develop an estimator for the OOS Sharpe ratio. The corresponding asymptotic properties of the Sharpe ratio estimator are established based on random matrix theory. Simulation results show that the proposed estimators perform close to the corresponding oracle level. Moreover, we numerically investigate the impact of various forms of regularization matrices on the OOS SR, which provides useful guidance for practical implementation. Finally, based on OOS SR estimator, we propose an adaptive regularized portfolio which uses the best regularization matrix yielding the highest estimated SR (among a set of candidates) at each decision node. Empirical evidence based on the S\&P 500 index demonstrates that the proposed adaptive ESG-constrained portfolio achieves a high OOS SR while satisfying the required ESG level, offering a practically effective approach for sustainable investment.

q-fin.PM

Equivalence Test for Mean Functions from Multi-population Functional Data

Most existing methods for testing equality of means of functional data from multiple populations rely on assumptions of equal covariance and/or Gaussianity. In this work we provide a new testing method based on a statistic that is distribution-free under the null hypothesis (i.e. the statistic is pivotal), and allows different covariance structures across populations, while Gaussianity is not required. In contrast to classical methods of functional mean testing, where either observations of the full curves or projections are applied, our method allows the projection dimension to increase with the sample size to allow asymptotic recovery of full information as the sample size increases. We obtain a unified theory for the asymptotic distribution of the test statistic under local alternatives, in both the sample and bootstrap cases. The finite sample performance for both size and power have been studied via simulations and the approach has also been applied to two real datasets.

stat.ME

Adaptive Multi-task Learning for Multi-sector Portfolio Optimization

Accurate transfer of information across multiple sectors to enhance model estimation is both significant and challenging in multi-sector portfolio optimization involving a large number of assets in different classes. Within the framework of factor modeling, we propose a novel data-adaptive multi-task learning methodology that quantifies and learns the relatedness among the principal temporal subspaces (spanned by factors) across multiple sectors under study. This approach not only improves the simultaneous estimation of multiple factor models but also enhances multi-sector portfolio optimization, which heavily depends on the accurate recovery of these factor models. Additionally, a novel and easy-to-implement algorithm, termed projection-penalized principal component analysis, is developed to accomplish the multi-task learning procedure. Diverse simulation designs and practical application on daily return data from Russell 3000 index demonstrate the advantages of multi-task learning methodology.

stat.ME

A Robust Extrinsic Single-index Model for Spherical Data

Regression with a spherical response is challenging due to the absence of linear structure, making standard regression models inadequate. Existing methods, mainly parametric, lack the flexibility to capture the complex relationship induced by spherical curvature, while methods based on techniques from Riemannian geometry often suffer from computational difficulties. The non-Euclidean structure further complicates robust estimation, with very limited work addressing this issue, despite the common presence of outliers in directional data. This article introduces a new semi-parametric approach, the extrinsic single-index model (ESIM) and its robust estimation, to address these limitations. We establish large-sample properties of the proposed estimator with a wide range of loss functions and assess their robustness using the influence function and standardized influence function. Specifically, we focus on the robustness of the exponential squared loss (ESL), demonstrating comparable efficiency and superior robustness over least squares loss under high concentration. We also examine how the tuning parameter for the ESL balances efficiency and robustness, providing guidance on its optimal choice. The computational efficiency and robustness of our methods are further illustrated via simulations and applications to geochemical compositional data.

stat.ME

Learning Fair Decisions with Factor Models: Applications to Annuity Pricing

Fairness-aware statistical learning is essential for mitigating discrimination against protected attributes such as gender, race, and ethnicity in data-driven decision-making. This is particularly critical in high-stakes applications like insurance underwriting and annuity pricing, where biased business decisions can have significant financial and social consequences. Factor models are commonly used in these domains for risk assessment and pricing; however, their predictive outputs may inadvertently introduce or amplify bias. To address this, we propose a Fair Decision Model that incorporates fairness regularization to mitigate outcome disparities. Specifically, the model is designed to ensure that expected decision errors are balanced across demographic groups - a criterion we refer to as Decision Error Parity. We apply this framework to annuity pricing based on mortality modelling. An empirical analysis using Australian mortality data demonstrates that the Fair Decision Model can significantly reduce decision error disparity while also improving predictive accuracy compared to benchmark models, including both traditional and fair factor models.

stat.ME

Double Descent in Portfolio Optimization: Dance between Theoretical Sharpe Ratio and Estimation Accuracy

We study the relationship between model complexity and out-of-sample performance in the context of mean-variance portfolio optimization. Representing model complexity by the number of assets, we find that the performance of low-dimensional models initially improves with complexity but then declines due to overfitting. As model complexity becomes sufficiently high, the performance improves with complexity again, resulting in a double ascent Sharpe ratio curve similar to the double descent phenomenon observed in artificial intelligence. The underlying mechanisms involve an intricate interaction between the theoretical Sharpe ratio and estimation accuracy. In high-dimensional models, the theoretical Sharpe ratio approaches its upper limit, and the overfitting problem is reduced because there are more parameters than data restrictions, which allows us to choose well-behaved parameters based on inductive bias.

q-fin.PM

Shocks-adaptive Robust Minimum Variance Portfolio for a Large Universe of Assets

This paper proposes a robust, shocks-adaptive portfolio in a large-dimensional assets universe where the number of assets could be comparable to or even larger than the sample size. It is well documented that portfolios based on optimizations are sensitive to outliers in return data. We deal with outliers by proposing a robust factor model, contributing methodologically through the development of a robust principal component analysis (PCA) for factor model estimation and a shrinkage estimation for the random error covariance matrix. This approach extends the well-regarded Principal Orthogonal Complement Thresholding (POET) method (Fan et al., 2013), enabling it to effectively handle heavy tails and sudden shocks in data. The novelty of the proposed robust method is its adaptiveness to both global and idiosyncratic shocks, without the need to distinguish them, which is useful in forming portfolio weights when facing outliers. We develop the theoretical results of the robust factor model and the robust minimum variance portfolio. Numerical and empirical results show the superior performance of the new portfolio.

q-fin.PM

Uncertainty Learning for High-dimensional Mean-variance Portfolio

Robust estimation for modern portfolio selection on a large set of assets becomes more important due to large deviation of empirical inference on big data. We propose a distributionally robust methodology for high-dimensional mean-variance portfolio problem, aiming to select an optimal conservative portfolio allocation by taking distribution uncertainty into account. With the help of factor structure, we extend the distributionally robust mean-variance problem investigated by Blanchet et al. (2022, Management Science) to the high-dimensional scenario and transform it to a new penalized risk minimization problem. Furthermore, we propose a data-adaptive method to estimate the quantified uncertainty size, which is the radius around the empirical probability measured by the Wasserstein distance. Asymptotic consistency is derived for the estimation of the population parameters involved in selecting the uncertainty size and the selected portfolio return. Our Monte-Carlo simulation results show that the chosen uncertainty size and target return from the proposed procedure are very close to the corresponding oracle version, and the new portfolio strategy is of low risk. Finally, we conduct empirical studies based on S&P index components to show the robust performance of our proposal in terms of risk controlling and return-risk balancing.

stat.ME

Localized Neural Network Modelling of Time Series: A Case Study on US Monetary Policy

In this paper, we investigate a semiparametric regression model under the context of treatment effects via a localized neural network (LNN) approach. Due to a vast number of parameters involved, we reduce the number of effective parameters by (i) exploring the use of identification restrictions; and (ii) adopting a variable selection method based on the group-LASSO technique. Subsequently, we derive the corresponding estimation theory and propose a dependent wild bootstrap procedure to construct valid inferences accounting for the dependence of data. Finally, we validate our theoretical findings through extensive numerical studies. In an empirical study, we revisit the impacts of a tightening monetary policy action on a variety of economic variables, including short-/long-term interest rate, inflation, unemployment rate, industrial price and equity return via the newly proposed framework using a monthly dataset of the US.

econ.EM

Spiked eigenvalues of high-dimensional sample autocovariance matrices: CLT and applications

High-dimensional autocovariance matrices play an important role in dimension reduction for high-dimensional time series. In this article, we establish the central limit theorem (CLT) for spiked eigenvalues of high-dimensional sample autocovariance matrices, which are developed under general conditions. The spiked eigenvalues are allowed to go to infinity in a flexible way without restrictions in divergence order. Moreover, the number of spiked eigenvalues and the time lag of the autocovariance matrix under this study could be either fixed or tending to infinity when the dimension p and the time length T go to infinity together. As a further statistical application, a novel autocovariance test is proposed to detect the equivalence of spiked eigenvalues for two high-dimensional time series. Various simulation studies are illustrated to justify the theoretical findings. Furthermore, a hierarchical clustering approach based on the autocovariance test is constructed and applied to clustering mortality data from multiple countries.

math.ST

Forecasting high-dimensional functional time series with dual-factor structures

We propose a dual-factor model for high-dimensional functional time series (HDFTS) that considers multiple populations. The HDFTS is first decomposed into a collection of functional time series (FTS) in a lower dimension and a group of population-specific basis functions. The system of basis functions describes cross-sectional heterogeneity, while the reduced-dimension FTS retains most of the information common to multiple populations. The low-dimensional FTS is further decomposed into a product of common functional loadings and a matrix-valued time series that contains the most temporal dynamics embedded in the original HDFTS. The proposed general-form dual-factor structure is connected to several commonly used functional factor models. We demonstrate the finite-sample performances of the proposed method in recovering cross-sectional basis functions and extracting common features using simulated HDFTS. An empirical study shows that the proposed model produces more accurate point and interval forecasts for subnational age-specific mortality rates in Japan. The financial benefits associated with the improved mortality forecasts are translated into a life annuity pricing scheme.

stat.ME

Robust PCA for High Dimensional Data based on Characteristic Transformation

In this paper, we propose a novel robust Principal Component Analysis (PCA) for high-dimensional data in the presence of various heterogeneities, especially the heavy-tailedness and outliers. A transformation motivated by the characteristic function is constructed to improve the robustness of the classical PCA. Besides the typical outliers, the proposed method has the unique advantage of dealing with heavy-tail-distributed data, whose covariances could be nonexistent (positively infinite, for instance). The proposed approach is also a case of kernel principal component analysis (KPCA) method and adopts the robust and non-linear properties via a bounded and non-linear kernel function. The merits of the new method are illustrated by some statistical properties including the upper bound of the excess error and the behaviors of the large eigenvalues under a spiked covariance model. In addition, we show the advantages of our method over the classical PCA by a variety of simulations. At last, we apply the new robust PCA to classify mice with different genotypes in a biological study based on their protein expression data and find that our method is more accurately on identifying abnormal mice comparing to the classical PCA.

stat.ME

Homogeneity and Sub-homogeneity Pursuit: Iterative Complement Clustering PCA

Principal component analysis (PCA), the most popular dimension-reduction technique, has been used to analyze high-dimensional data in many areas. It discovers the homogeneity within the data and creates a reduced feature space to capture as much information as possible from the original data. However, in the presence of a group structure of the data, PCA often fails to identify the group-specific pattern, which is known as sub-homogeneity in this study. Group-specific information that is missed can result in an unsatisfactory representation of the data from a particular group. It is important to capture both homogeneity and sub-homogeneity in high-dimensional data analysis, but this poses a great challenge. In this study, we propose a novel iterative complement-clustering principal component analysis (CPCA) to iteratively estimate the homogeneity and sub-homogeneity. A principal component regression based clustering method is also introduced to provide reliable information about clusters. Theoretically, this study shows that our proposed clustering approach can correctly identify the cluster membership under certain conditions. The simulation study and real analysis of the stock return data confirm the superior performance of our proposed methods.

stat.ME

Factor-augmented model for functional data

We propose modeling raw functional data as a mixture of a smooth function and a high-dimensional factor component. The conventional approach to retrieving the smooth function from the raw data is through various smoothing techniques. However, the smoothing model is inadequate to recover the smooth curve or capture the data variation in some situations. These include cases where there is a large amount of measurement error, the smoothing basis functions are incorrectly identified, or the step jumps in the functional mean levels are neglected. A factor-augmented smoothing model is proposed to address these challenges, and an iterative numerical estimation approach is implemented in practice. Including the factor model component in the proposed method solves the aforementioned problems since a few common factors often drive the variation that cannot be captured by the smoothing model. Asymptotic theorems are also established to demonstrate the effects of including factor structures on the smoothing results. Specifically, we show that the smoothing coefficients projected on the complement space of the factor loading matrix are asymptotically normal. As a byproduct of independent interest, an estimator for the population covariance matrix of the raw data is presented based on the proposed model. Extensive simulation studies illustrate that these factor adjustments are essential in improving estimation accuracy and avoiding the curse of dimensionality. The superiority of our model is also shown in modeling Australian temperature data.

stat.ME

Clustering and Forecasting Multiple Functional Time Series

Modelling and forecasting homogeneous age-specific mortality rates of multiple countries could lead to improvements in long-term forecasting. Data fed into joint models are often grouped according to nominal attributes, such as geographic regions, ethnic groups, and socioeconomic status, which may still contain heterogeneity and deteriorate the forecast results. Our paper proposes a novel clustering technique to pursue homogeneity among multiple functional time series based on functional panel data modelling to address this issue. Using a functional panel data model with fixed effects, we can extract common functional time series features. These common features could be decomposed into two components: the functional time trend and the mode of variations of functions (functional pattern). The functional time trend reflects the dynamics across time, while the functional pattern captures the fluctuations within curves. The proposed clustering method searches for homogeneous age-specific mortality rates of multiple countries by accounting for both the modes of variations and the temporal dynamics among curves. We demonstrate that the proposed clustering technique outperforms other existing methods through a Monte Carlo simulation and could handle complicated cases with slow decaying eigenvalues. In empirical data analysis, we find that the clustering results of age-specific mortality rates can be explained by the combination of geographic region, ethnic groups, and socioeconomic status. We further show that our model produces more accurate forecasts than several benchmark methods in forecasting age-specific mortality rates.

stat.ME

A Forecast-driven Hierarchical Factor Model with Application to Mortality Data

Mortality forecasting plays a pivotal role in insurance and financial risk management of life insurers, pension funds, and social securities. Mortality data is usually high-dimensional in nature and favors factor model approaches to modelling and forecasting. This paper introduces a new forecast-driven hierarchical factor model (FHFM) customized for mortality forecasting. Compared to existing models, which only capture the cross-sectional variation or time-serial dependence in the dimension reduction step, the new model captures both features efficiently under a hierarchical structure, and provides insights into the understanding of dynamic variation of mortality patterns over time. By comparing with static PCA utilized in Lee and Carter 1992, dynamic PCA introduced in Lam et al. 2011, as well as other existing mortality modelling methods, we find that this approach provides both better estimation results and superior out-of-sample forecasting performance. Simulation studies further illustrate the advantages of the proposed model based on different data structures. Finally, empirical studies using the US mortality data demonstrate the implications and significance of this new model in life expectancy forecasting and life annuities pricing.

stat.AP