Searcharxiv⌕ Search

arXiv subjects

Han Lin Shang

Publications and source records attributed to Han Lin Shang.

At least 109 records · Page 6Linked to original sources

On function-on-function regression: Partial least squares approach

Functional data analysis tools, such as function-on-function regression models, have received considerable attention in various scientific fields because of their observed high-dimensional and complex data structures. Several statistical procedures, including least squares, maximum likelihood, and maximum penalized likelihood, have been proposed to estimate such function-on-function regression models. However, these estimation techniques produce unstable estimates in the case of degenerate functional data or are computationally intensive. To overcome these issues, we proposed a partial least squares approach to estimate the model parameters in the function-on-function regression model. In the proposed method, the B-spline basis functions are utilized to convert discretely observed data into their functional forms. Generalized cross-validation is used to control the degrees of roughness. The finite-sample performance of the proposed method was evaluated using several Monte-Carlo simulations and an empirical data analysis. The results reveal that the proposed method competes favorably with existing estimation techniques and some other available function-on-function regression models, with significantly shorter computational time.

stat.ME↗

Implied volatility surface predictability: the case of commodity markets

Recent literature seek to forecast implied volatility derived from equity, index, foreign exchange, and interest rate options using latent factor and parametric frameworks. Motivated by increased public attention borne out of the financialization of futures markets in the early 2000s, we investigate if these extant models can uncover predictable patterns in the implied volatility surfaces of the most actively traded commodity options between 2006 and 2016. Adopting a rolling out-of-sample forecasting framework that addresses the common multiple comparisons problem, we establish that, for energy and precious metals options, explicitly modeling the term structure of implied volatility using the Nelson-Siegel factors produces the most accurate forecasts.

q-fin.ST↗

Dynamic principal component regression for forecasting functional time series in a group structure

When generating social policies and pricing annuity at national and subnational levels, it is essential both to forecast mortality accurately and ensure that forecasts at the subnational level add up to the forecasts at the national level. This has motivated recent developments in forecasting functional time series in a group structure, where static principal component analysis is used. In the presence of moderate to strong temporal dependence, static principal component analysis designed for independent and identically distributed functional data may be inadequate. Thus, through using the dynamic functional principal component analysis, we consider a functional time series forecasting method with static and dynamic principal component regression to forecast each series in a group structure. Through using the regional age-specific mortality rates in Japan obtained from the Japanese Mortality Database (2019), we investigate the point and interval forecast accuracies of our proposed extension, and subsequently make recommendations.

stat.AP↗

Forecasting age distribution of death counts: An application to annuity pricing

We consider a compositional data analysis approach to forecasting the age distribution of death counts. Using the age-specific period life-table death counts in Australia obtained from the Human Mortality Database, the compositional data analysis approach produces more accurate one- to 20-step-ahead point and interval forecasts than Lee-Carter method, Hyndman-Ullah method, and two naïve random walk methods. The improved forecast accuracy of period life-table death counts is of great interest to demographers for estimating survival probabilities and life expectancy, and to actuaries for determining temporary annuity prices for various ages and maturities. Although we focus on temporary annuity prices, we consider long-term contracts which make the annuity almost lifetime, in particular when the age at entry is sufficiently high.

stat.AP↗

Forecasting functional time series using weighted likelihood methodology

Functional time series whose sample elements are recorded sequentially over time are frequently encountered with increasing technology. Recent studies have shown that analyzing and forecasting of functional time series can be performed easily using functional principal component analysis and existing univariate/multivariate time series models. However, the forecasting performance of such functional time series models may be affected by the presence of outlying observations which are very common in many scientific fields. Outliers may distort the functional time series model structure, and thus, the underlying model may produce high forecast errors. We introduce a robust forecasting technique based on weighted likelihood methodology to obtain point and interval forecasts in functional time series in the presence of outliers. The finite sample performance of the proposed method is illustrated by Monte Carlo simulations and four real-data examples. Numerical results reveal that the proposed method exhibits superior performance compared with the existing method(s).

stat.ME↗

Dynamic principal component regression: Application to age-specific mortality forecasting

In areas of application, including actuarial science and demography, it is increasingly common to consider a time series of curves; an example of this is age-specific mortality rates observed over a period of years. Given that age can be treated as a discrete or continuous variable, a dimension reduction technique, such as principal component analysis, is often implemented. However, in the presence of moderate to strong temporal dependence, static principal component analysis commonly used for analyzing independent and identically distributed data may not be adequate. As an alternative, we consider a \textit{dynamic} principal component approach to model temporal dependence in a time series of curves. Inspired by Brillinger's (1974) theory of dynamic principal components, we introduce a dynamic principal component analysis, which is based on eigen-decomposition of estimated long-run covariance. Through a series of empirical applications, we demonstrate the potential improvement of one-year-ahead point and interval forecast accuracies that the dynamic principal component regression entails when compared with the static counterpart.

stat.AP↗

A robust functional time series forecasting method

Univariate time series often take the form of a collection of curves observed sequentially over time. Examples of these include hourly ground-level ozone concentration curves. These curves can be viewed as a time series of functions observed at equally spaced intervals over a dense grid. Since functional time series may contain various types of outliers, we introduce a robust functional time series forecasting method to down-weigh the influence of outliers in forecasting. Through a robust principal component analysis based on projection pursuit, a time series of functions can be decomposed into a set of robust dynamic functional principal components and their associated scores. Conditioning on the estimated functional principal components, the crux of the curve-forecasting problem lies in modeling and forecasting principal component scores, through a robust vector autoregressive forecasting method. Via a simulation study and an empirical study on forecasting ground-level ozone concentration, the robust method demonstrates the superior forecast accuracy that dynamic functional principal component regression entails. The robust method also shows the superior estimation accuracy of the parameters in the vector autoregressive models for modeling and forecasting principal component scores, and thus improves curve forecast accuracy.

stat.ME↗

Uncovering predictability in the evolution of the WTI oil futures curve

Accurately forecasting the price of oil, the world's most actively traded commodity, is of great importance to both academics and practitioners. We contribute by proposing a functional time series based method to model and forecast oil futures. Our approach boasts a number of theoretical and practical advantages including effectively exploiting underlying process dynamics missed by classical discrete approaches. We evaluate the finite-sample performance against established benchmarks using a model confidence set test. A realistic out-of-sample exercise provides strong support for the adoption of our approach with it residing in the superior set of models in all considered instances.

stat.AP↗

Intraday forecasts of a volatility index: Functional time series methods with dynamic updating

As a forward-looking measure of future equity market volatility, the VIX index has gained immense popularity in recent years to become a key measure of risk for market analysts and academics. We consider discrete reported intraday VIX tick values as realisations of a collection of curves observed sequentially on equally spaced and dense grids over time and utilise functional data analysis techniques to produce one-day-ahead forecasts of these curves. The proposed method facilitates the investigation of dynamic changes in the index over very short time intervals as showcased using the 15-second high-frequency VIX index values. With the help of dynamic updating techniques, our point and interval forecasts are shown to enjoy improved accuracy over conventional time series models.

stat.AP↗

Estimation of a functional single index model with dependent errors and unknown error density

The problem of error density estimation for a functional single index model with dependent errors is studied. A Bayesian method is utilized to simultaneously estimate the bandwidths in the kernel-form error density and regression function, under an autoregressive error structure. For estimating both the regression function and error density, empirical studies show that the functional single index model gives improved estimation and prediction accuracies than any nonparametric functional regression considered. Furthermore, estimation of error density facilitates the construction of prediction interval for the response variable.

stat.AP↗

Semiparametric Regression using Variational Approximations

Semiparametric regression offers a flexible framework for modeling non-linear relationships between a response and covariates. A prime example are generalized additive models where splines (say) are used to approximate non-linear functional components in conjunction with a quadratic penalty to control for overfitting. Estimation and inference are then generally performed based on the penalized likelihood, or under a mixed model framework. The penalized likelihood framework is fast but potentially unstable, and choosing the smoothing parameters needs to be done externally using cross-validation, for instance. The mixed model framework tends to be more stable and offers a natural way for choosing the smoothing parameters, but for non-normal responses involves an intractable integral. In this article, we introduce a new framework for semiparametric regression based on variational approximations. The approach possesses the stability and natural inference tools of the mixed model framework, while achieving computation times comparable to using penalized likelihood. Focusing on generalized additive models, we derive fully tractable variational likelihoods for some common response types. We present several features of the variational approximation framework for inference, including a variational information matrix for inference on parametric components, and a closed-form update for estimating the smoothing parameter. We demonstrate the consistency of the variational approximation estimates, and an asymptotic normality result for the parametric component of the model. Simulation studies show the variational approximation framework performs similarly to and sometimes better than currently available software for fitting generalized additive models.

math.ST↗

High-dimensional functional time series forecasting: An application to age-specific mortality rates

We address the problem of forecasting high-dimensional functional time series through a two-fold dimension reduction procedure. The difficulty of forecasting high-dimensional functional time series lies in the curse of dimensionality. In this paper, we propose a novel method to solve this problem. Dynamic functional principal component analysis is first applied to reduce each functional time series to a vector. We then use the factor model as a further dimension reduction technique so that only a small number of latent factors are preserved. Classic time series models can be used to forecast the factors and conditional forecasts of the functions can be constructed. Asymptotic properties of the approximated functions are established, including both estimation error and forecast error. The proposed method is easy to implement especially when the dimension of the functional time series is large. We show the superiority of our approach by both simulation studies and an application to Japanese age-specific mortality rates.

stat.ME↗

Model confidence sets and forecast combination: An application to age-specific mortality

Model averaging combines forecasts obtained from a range of models, and it often produces more accurate forecasts than a forecast from a single model. The crucial part of forecast accuracy improvement in using the model averaging lies in the determination of optimal weights from a finite sample. If the weights are selected sub-optimally, this can affect the accuracy of the model-averaged forecasts. Instead of choosing the optimal weights, we consider trimming a set of models before equally averaging forecasts from the selected superior models. Motivated by Hansen, Lunde and Nason (2011), we apply and evaluate the model confidence set procedure when combining mortality forecasts. The proposed model averaging procedure is motivated by Samuels and Sekkel (2017) based on the concept of model confidence sets as proposed by Hansen et al. (2011) that incorporates the statistical significance of the forecasting performance. As the model confidence level increases, the set of superior models generally decreases. The proposed model averaging procedure is demonstrated via national and sub-national Japanese mortality for retirement ages between 60 and 100+. Illustrated by national and sub-national Japanese mortality for ages between 60 and 100+, the proposed model-average procedure gives the smallest interval forecast errors, especially for males. We find that robust out-of-sample point and interval forecasts may be obtained from the trimming method. By robust, we mean robustness against model misspecification.

stat.AP↗

Visualising rate of change: application to age-specific fertility

Visualisation methods help in the discovery of characteristics that might not have been apparent using mathematical models and summary statistics. However, visualisation methods have not received much attention in demography, with the exceptions of scatter plot and Lexis surface. We utilise a phase-plane plot to visualise the rate of change, obtained from derivatives of a continuous function. The phase-plane plot bears a resemblance to hysteresis loops, isogrowth curves, and solutions to differential equations. Using Australian and Chilean fertility, we present phase-plane plots. Similarly to the scatter plot and Lexis surface, the phase-plane plot identifies the age with maximum fertility rate and displays skewness of fertility distribution. Unlike the scatter plot and Lexis surface, the phase-plane plot identifies the age with maximum positive or negative velocity (i.e., trend), can compare the magnitude of the rate of change between any two years based on the size of the radius of circles. The phase-plane plot allows the visualisation of dynamic changes in fertility for a given age over the years and is potentially useful for visualising dynamic changes in birth-cohort fertility. Via the animate package in LaTeX, a dynamic phase-plane plot is also proposed to visualise changes in fertility over age or year.

stat.AP↗

Grouped multivariate and functional time series forecasting: An application to annuity pricing

Age-specific mortality rates are often disaggregated by different attributes, such as sex, state, ethnic group and socioeconomic status. In making social policies and pricing annuity at national and subnational levels, it is important not only to forecast mortality accurately, but also to ensure that forecasts at the subnational level add up to the forecasts at the national level. This motivates recent developments in grouped functional time series methods (Shang and Hyndman, 2017) to reconcile age-specific mortality forecasts. We extend these grouped functional time series forecasting methods to multivariate time series, and apply them to produce point forecasts of mortality rates at older ages, from which fixed-term annuities for different ages and maturities can be priced. Using the regional age-specific mortality rates in Japan obtained from the Japanese Mortality Database, we investigate the one-step-ahead to 15-step-ahead point-forecast accuracy between the independent and grouped forecasting methods. The grouped forecasting methods are shown not only to be useful for reconciling forecasts of age-specific mortality rates at national and subnational levels, but they are also shown to allow improved forecast accuracy. The improved forecast accuracy of mortality rates is of great interest to the insurance and pension industries for estimating annuity prices, in particular at the level of population subgroups, defined by key factors such as sex, region, and socioeconomic grouping.

stat.AP↗

Bootstrap methods for stationary functional time series

Bootstrap methods for estimating the long-run covariance of stationary functional time series are considered. We introduce a versatile bootstrap method that relies on functional principal component analysis, where principal component scores can be bootstrapped by maximum entropy. Two other bootstrap methods resample error functions, after the dependence structure being modeled linearly by a sieve method or nonlinearly by a functional kernel regression. Through a series of Monte-Carlo simulation, we evaluate and compare the finite-sample performances of these three bootstrap methods for estimating the long-run covariance in a functional time series. Using the intraday particulate matter (PM10) data set in Graz, the proposed bootstrap methods provide a way of constructing the distribution of estimated long-run covariance for functional time series.

stat.CO↗

Mortality and life expectancy forecasting for a group of populations in developed countries: A robust multilevel functional data method

A robust multilevel functional data method is proposed to forecast age-specific mortality rate and life expectancy for two or more populations in developed countries with high-quality vital registration systems. It uses a robust multilevel functional principal component analysis of aggregate and population-specific data to extract the common trend and population-specific residual trend among populations. This method is applied to age- and sex-specific mortality rate and life expectancy for the United Kingdom from 1922 to 2011, and its forecast accuracy is then further compared with standard multilevel functional data method. For forecasting both age-specific mortality and life expectancy, the robust multilevel functional data method produces more accurate point and interval forecasts than the standard multilevel functional data method in the presence of outliers.

stat.AP↗

Grouped functional time series forecasting: An application to age-specific mortality rates

Age-specific mortality rates are often disaggregated by different attributes, such as sex, state and ethnicity. Forecasting age-specific mortality rates at the national and sub-national levels plays an important role in developing social policy. However, independent forecasts at the sub-national levels may not add up to the forecasts at the national level. To address this issue, we consider reconciling forecasts of age-specific mortality rates, extending the methods of Hyndman et al. (2011) to functional time series, where age is considered as a continuum. The grouped functional time series methods are used to produce point forecasts of mortality rates that are aggregated appropriately across different disaggregation factors. For evaluating forecast uncertainty, we propose a bootstrap method for reconciling interval forecasts. Using the regional age-specific mortality rates in Japan, obtained from the Japanese Mortality Database, we investigate the one- to ten-step-ahead point and interval forecast accuracies between the independent and grouped functional time series forecasting methods. The proposed methods are shown to be useful for reconciling forecasts of age-specific mortality rates at the national and sub-national levels. They also enjoy improved forecast accuracy averaged over different disaggregation factors. Supplemental materials for the article are available online.

stat.AP↗