Searcharxiv⌕ Search

arXiv subjects

Han Lin Shang

Publications and source records attributed to Han Lin Shang.

At least 55 records · Page 3Linked to original sources

Constructing prediction intervals for the age distribution of deaths

We introduce a model-agnostic procedure to construct prediction intervals for the age distribution of deaths. The age distribution of deaths is an example of constrained data, which are nonnegative and have a constrained integral. A centered log-ratio transformation and a cumulative distribution function transformation are used to remove the two constraints, where the latter transformation can also handle the presence of zero counts. Our general procedure divides data samples into training, validation, and testing sets. Within the validation set, we can select an optimal tuning parameter by calibrating the empirical coverage probabilities to be close to their nominal ones. With the selected optimal tuning parameter, we then construct the pointwise prediction intervals using the same models for the holdout data in the testing set. Using Japanese age- and sex-specific life-table death counts, we assess and evaluate the interval forecast accuracy with a suite of functional time-series models.

stat.ME↗

Bootstrap prediction intervals for the age distribution of life-table death counts

We introduce a nonparametric bootstrap procedure based on a dynamic factor model to construct pointwise prediction intervals for period life-table death counts. The age distribution of death counts is an example of constrained data, which are nonnegative and have a constrained integral. A centered log-ratio transformation is used to remove the constraints. With a time series of unconstrained data, we introduce our bootstrap method to construct prediction intervals, thereby quantifying forecast uncertainty. The bootstrap method utilizes a dynamic factor model to capture both nonstationary and stationary patterns through a two-stage functional principal component analysis. To capture parameter uncertainty, the estimated principal component scores and model residuals are sampled with replacement. Using the age- and sex-specific life-table deaths for Australia and the United Kingdom, we study the empirical coverage probabilities and compare them with the nominal ones. The bootstrap method has superior interval forecast accuracy, especially for the one-step-ahead forecast horizon.

stat.ME↗

Spatial Functional Deep Neural Network Model: A New Prediction Algorithm

Accurate prediction of spatially dependent functional data is critical for various engineering and scientific applications. In this study, a spatial functional deep neural network model was developed with a novel non-linear modeling framework that seamlessly integrates spatial dependencies and functional predictors using deep learning techniques. The proposed model extends classical scalar-on-function regression by incorporating a spatial autoregressive component while leveraging functional deep neural networks to capture complex non-linear relationships. To ensure a robust estimation, the methodology employs an adaptive estimation approach, where the spatial dependence parameter was first inferred via maximum likelihood estimation, followed by non-linear functional regression using deep learning. The effectiveness of the proposed model was evaluated through extensive Monte Carlo simulations and an application to Brazilian COVID-19 data, where the goal was to predict the average daily number of deaths. Comparative analysis with maximum likelihood-based spatial functional linear regression and functional deep neural network models demonstrates that the proposed algorithm significantly improves predictive performance. The results for the Brazilian COVID-19 data showed that while all models achieved similar mean squared error values over the training modeling phase, the proposed model achieved the lowest mean squared prediction error in the testing phase, indicating superior generalization ability.

stat.ME↗

Forecasting a time series of Lorenz curves: One-way functional analysis of variance

The Lorenz curve is a fundamental tool for analysing income and wealth distribution and inequality at national and regional levels. We utilise a one-way functional analysis of variance to decompose a time series of Lorenz curves and develop a method for producing one-step-ahead point and interval forecasts. The one-way functional analysis of variance is easily interpretable by decomposing an array into a functional grand effect, a functional row effect and residual functions. We evaluate and compare the forecast accuracy between the functional analysis of variance and three non-functional methods using the Italian household income and wealth data.

stat.ME↗

Dependence-based fuzzy clustering of functional time series

Time series clustering is essential in scientific applications, yet methods for functional time series, collections of infinite-dimensional curves treated as random elements in a Hilbert space, remain underdeveloped. This work presents clustering approaches for functional time series that combine the fuzzy $C$-medoids and fuzzy $C$-means procedures with a novel dissimilarity measure tailored for functional data. This dissimilarity is based on an extension of the quantile autocorrelation to the functional context. Our methods effectively groups time series with similar dependence structures, achieving high accuracy and computational efficiency in simulations. The practical utility of the approach is demonstrated through case studies on high-frequency financial stock data and multi-country age-specific mortality improvements.

stat.ME↗

Forecasting Age Distribution of Deaths: Cumulative Distribution Function Transformation

Like density functions, period life-table death counts are nonnegative and have a constrained integral, and thus live in a constrained nonlinear space. Implementing established modelling and forecasting methods without obeying these constraints can be problematic for such nonlinear data. We introduce cumulative distribution function transformation to forecast the life-table death counts. Using the Japanese life-table death counts obtained from the Japanese Mortality Database (2024), we evaluate the point and interval forecast accuracies of the proposed approach, which compares favourably to an existing compositional data analytic approach. The improved forecast accuracy of life-table death counts is of great interest to demographers for estimating age-specific survival probabilities and life expectancy and actuaries for determining temporary annuity prices for different ages and maturities.

stat.ME↗

Stock Return Prediction based on a Functional Capital Asset Pricing Model

The capital asset pricing model (CAPM) is readily used to capture a linear relationship between the daily returns of an asset and a market index. We extend this model to an intraday high-frequency setting by proposing a functional CAPM estimation approach. The functional CAPM is a stylized example of a function-on-function linear regression with a bivariate functional regression coefficient. The two-dimensional regression coefficient measures the cross-covariance between cumulative intraday asset returns and market returns. We apply it to the Standard and Poor's 500 index and its constituent stocks to demonstrate its practicality. We investigate the functional CAPM's in-sample goodness-of-fit and out-of-sample prediction for an asset's cumulative intraday return. The findings suggest that the proposed functional CAPM methods have superior model goodness-of-fit and forecast accuracy compared to the traditional CAPM empirical estimation. In particular, the functional methods produce better model goodness-of-fit and prediction accuracy for stocks traditionally considered less price-efficient or more information-opaque.

stat.ME↗

Penalized function-on-function linear quantile regression

We introduce a novel function-on-function linear quantile regression model to characterize the entire conditional distribution of a functional response for a given functional predictor. Tensor cubic $B$-splines expansion is used to represent the regression parameter functions, where a derivative-free optimization algorithm is used to obtain the estimates. Quadratic roughness penalties are applied to the coefficients to control the smoothness of the estimates. The optimal degree of smoothness depends on the quantile of interest. An automatic grid-search algorithm based on the Bayesian information criterion is used to estimate the optimum values of the smoothing parameters. Via a series of Monte-Carlo experiments and an empirical data analysis using Mary River flow data, we evaluate the estimation and predictive performance of the proposed method, and the results are compared favorably with several existing methods.

stat.ME↗

Forecasting density-valued functional panel data

We introduce a statistical method for modeling and forecasting functional panel data represented by multiple densities. Density functions are nonnegative and have a constrained integral and thus do not constitute a linear vector space. We implement a center log-ratio transformation to transform densities into unconstrained functions. These functions exhibit cross-sectional correlation and temporal dependence. Via a functional analysis of variance decomposition, we decompose the unconstrained functional panel data into a deterministic trend component and a time-varying residual component. To produce forecasts for the time-varying component, a functional time series forecasting method, based on the estimation of the long-run covariance, is implemented. By combining the forecasts of the time-varying residual component with the deterministic trend component, we obtain $h$-step-ahead forecast curves for multiple populations. Illustrated by age- and sex-specific life-table death counts in the United States, we apply our proposed method to generate forecasts of the life-table death counts for 51 states.

stat.ME↗

Spatial function-on-function regression

We introduce a spatial function-on-function regression model to capture spatial dependencies in functional data by integrating spatial autoregressive techniques with functional principal component analysis. The proposed model addresses a critical gap in functional regression by enabling the analysis of functional responses influenced by spatially correlated functional predictors, a common scenario in fields such as environmental sciences, epidemiology, and socio-economic studies. The model employs a spatial functional principal component decomposition on the response and a classical functional principal component decomposition on the predictor, transforming the functional data into a finite-dimensional multivariate spatial autoregressive framework. This transformation allows efficient estimation and robust handling of spatial dependencies through least squares methods. In a series of extensive simulations, the proposed model consistently demonstrated superior performance in estimating both spatial autocorrelation and regression coefficient functions compared to some favorably existing traditional approaches, particularly under moderate to strong spatial effects. Application of the proposed model to Brazilian COVID-19 data further underscored its practical utility, revealing critical spatial patterns in confirmed cases and death rates that align with known geographic and social interactions. An R package provides a comprehensive implementation of the proposed estimation method, offering a user-friendly and efficient tool for researchers and practitioners to apply the methodology in real-world scenarios.

stat.ME↗

Nonstationary functional time series forecasting

We propose a nonstationary functional time series forecasting method with an application to age-specific mortality rates observed over the years. The method begins by taking the first-order differencing and estimates its long-run covariance function. Through eigen-decomposition, we obtain a set of estimated functional principal components and their associated scores for the differenced series. These components allow us to reconstruct the original functional data and compute the residuals. To model the temporal patterns in the residuals, we again perform dynamic functional principal component analysis and extract its estimated principal components and the associated scores for the residuals. As a byproduct, we introduce a geometrically decaying weighted approach to assign higher weights to the most recent data than those from the distant past. Using the Swedish age-specific mortality rates from 1751 to 2022, we demonstrate that the weighted dynamic functional factor model can produce more accurate point and interval forecasts, particularly for male series exhibiting higher volatility.

stat.ME↗

Change-point detection in functional time series: Applications to age-specific mortality and fertility

We consider determining change points in a time series of age-specific mortality and fertility curves observed over time. We propose two detection methods for identifying these change points. The first method uses a functional cumulative sum statistic to pinpoint the change point. The second method computes a univariate time series of integrated squared forecast errors after fitting a functional time-series model before applying a change-point detection method to the errors to determine the change point. Using Australian age-specific fertility and mortality data, we apply these methods to locate the change points and identify the optimal training period to achieve improved forecast accuracy.

stat.AP↗

Robust function-on-function interaction regression

A function-on-function regression model with quadratic and interaction effects of the covariates provides a more flexible model. Despite several attempts to estimate the model's parameters, almost all existing estimation strategies are non-robust against outliers. Outliers in the quadratic and interaction effects may deteriorate the model structure more severely than their effects in the main effect. We propose a robust estimation strategy based on the robust functional principal component decomposition of the function-valued variables and $τ$-estimator. The performance of the proposed method relies on the truncation parameters in the robust functional principal component decomposition of the function-valued variables. A robust Bayesian information criterion is used to determine the optimum truncation constants. A forward stepwise variable selection procedure is employed to determine relevant main, quadratic, and interaction effects to address a possible model misspecification. The finite-sample performance of the proposed method is investigated via a series of Monte-Carlo experiments. The proposed method's asymptotic consistency and influence function are also studied in the supplement, and its empirical performance is further investigated using a U.S. COVID-19 dataset.

stat.ME↗

Forecasting Australian fertility by age, region, and birthplace

Fertility differentials by urban-rural residence and nativity of women in Australia significantly impact population composition at sub-national levels. We aim to provide consistent fertility forecasts for Australian women characterized by age, region, and birthplace. Age-specific fertility rates at the national and sub-national levels obtained from census data between 1981-2011 are jointly modeled and forecast by the grouped functional time series method. Forecasts for women of each region and birthplace are reconciled following the chosen hierarchies to ensure that results at various disaggregation levels consistently sum up to the respective national total. Coupling the region of residence disaggregation structure with the trace minimization reconciliation method produces the most accurate point and interval forecasts. In addition, age-specific fertility rates disaggregated by the birthplace of women show significant heterogeneity that supports the application of the grouped forecasting method.

stat.AP↗

Enhancing Spatial Functional Linear Regression with Robust Dimension Reduction Methods

This paper introduces a robust estimation strategy for the spatial functional linear regression model using dimension reduction methods, specifically functional principal component analysis (FPCA) and functional partial least squares (FPLS). These techniques are designed to address challenges associated with spatially correlated functional data, particularly the impact of outliers on parameter estimation. By projecting the infinite-dimensional functional predictor onto a finite-dimensional space defined by orthonormal basis functions and employing M-estimation to mitigate outlier effects, our approach improves the accuracy and reliability of parameter estimates in the spatial functional linear regression context. Simulation studies and empirical data analysis substantiate the effectiveness of our methods, while an appendix explores the Fisher consistency and influence function of the FPCA-based approach. The rfsac package in R implements these robust estimation strategies, ensuring practical applicability for researchers and practitioners.

stat.ME↗

Is the age pension in Australia sustainable and fair? Evidence from forecasting the old-age dependency ratio using the Hamilton-Perry model

The age pension aims to assist eligible elderly Australians meet specific age and residency criteria in maintaining basic living standards. In designing efficient pension systems, government policymakers seek to satisfy the expectations of the overall aging population in Australia. However, the population's unique demographic characteristics at the state and territory level are often overlooked due to the lack of available data. We use the Hamilton-Perry model, which requires minimum input, to model and forecast the evolution of age-specific populations at the state level. We also integrate the obtained sub-national demographic information to determine sustainable pension ages up to 2051. We also investigate pension welfare distribution in all states and territories to identify disadvantaged residents under the current pension system. Using the sub-national mortality data for Australia from 1971 to 2021 obtained from AHMD (2023), we implement the Hamilton-Perry model with the help of functional time series forecasting techniques. With forecasts of age-specific population sizes for each state and territory, we compute the old age dependency ratio to determine the nationwide sustainable pension age.

stat.AP↗

Forecasting age distribution of life-table death counts via α-transformation

We introduce a compositional power transformation, known as an α-transformation, to model and forecast a time series of life-table death counts, possibly with zero counts observed at older ages. As a generalisation of the isometric log-ratio transformation (i.e., α = 0), the α transformation relies on the tuning parameter α, which can be determined in a data-driven manner. Using the Australian age-specific period life-table death counts from 1921 to 2020, the α transformation can produce more accurate short-term point and interval forecasts than the log-ratio transformation. The improved forecast accuracy of life-table death counts is of great importance to demographers and government planners for estimating survival probabilities and life expectancy and actuaries for determining annuity prices and reserves for various initial ages and maturity terms.

stat.AP↗

Uncertainty Learning for High-dimensional Mean-variance Portfolio

Robust estimation for modern portfolio selection on a large set of assets becomes more important due to large deviation of empirical inference on big data. We propose a distributionally robust methodology for high-dimensional mean-variance portfolio problem, aiming to select an optimal conservative portfolio allocation by taking distribution uncertainty into account. With the help of factor structure, we extend the distributionally robust mean-variance problem investigated by Blanchet et al. (2022, Management Science) to the high-dimensional scenario and transform it to a new penalized risk minimization problem. Furthermore, we propose a data-adaptive method to estimate the quantified uncertainty size, which is the radius around the empirical probability measured by the Wasserstein distance. Asymptotic consistency is derived for the estimation of the population parameters involved in selecting the uncertainty size and the selected portfolio return. Our Monte-Carlo simulation results show that the chosen uncertainty size and target return from the proposed procedure are very close to the corresponding oracle version, and the new portfolio strategy is of low risk. Finally, we conduct empirical studies based on S&P index components to show the robust performance of our proposal in terms of risk controlling and return-risk balancing.

stat.ME↗