Searcharxiv⌕ Search

arXiv subjects

Han Lin Shang

Publications and source records attributed to Han Lin Shang.

At least 91 records · Page 5Linked to original sources

Feature Extraction for Functional Time Series: Theory and Application to NIR Spectroscopy Data

We propose a novel method to extract global and local features of functional time series. The global features concerning the dominant modes of variation over the entire function domain, and local features of function variations over particular short intervals within function domain, are both important in functional data analysis. Functional principal component analysis (FPCA), though a key feature extraction tool, only focus on capturing the dominant global features, neglecting highly localized features. We introduce a FPCA-BTW method that initially extracts global features of functional data via FPCA, and then extracts local features by block thresholding of wavelet (BTW) coefficients. Using Monte Carlo simulations, along with an empirical application on near-infrared spectroscopy data of wood panels, we illustrate that the proposed method outperforms competing methods including FPCA and sparse FPCA in the estimation functional processes. Moreover, extracted local features inheriting serial dependence of the original functional time series contribute to more accurate forecasts. Finally, we develop asymptotic properties of FPCA-BTW estimators, discovering the interaction between convergence rates of global and local features.

stat.ME↗

A functional autoregressive model based on exogenous hydrometeorological variables for river flow prediction

In this research, a functional time series model was introduced to predict future realizations of river flow time series. The proposed model was constructed based on a functional time series's correlated lags and the essential exogenous climate variables. Rainfall, temperature, and evaporation variables were hypothesized to have substantial functionality in river flow simulation. Because an actual time series model is unspecified and the input variables' significance for the learning process is unknown in practice, it was employed a variable selection procedure to determine only the significant variables for the model. A nonparametric bootstrap model was also proposed to investigate predictions' uncertainty and construct pointwise prediction intervals for the river flow curve time series. Historical datasets at three meteorological stations (Mosul, Baghdad, and Kut) located in the semi-arid region, Iraq, were used for model development. The prediction performance of the proposed model was validated against existing functional and traditional time series models. The numerical analyses revealed that the proposed model provides competitive or even better performance than the benchmark models. Also, the incorporated exogenous climate variables have substantially improved the modeling predictability performance. Overall, the proposed model indicated a reliable methodology for modeling river flow within the semi-arid region.

stat.AP↗

Factor-augmented Smoothing Model for Functional Data

We propose modeling raw functional data as a mixture of a smooth function and a highdimensional factor component. The conventional approach to retrieving the smooth function from the raw data is through various smoothing techniques. However, the smoothing model is not adequate to recover the smooth curve or capture the data variation in some situations. These include cases where there is a large amount of measurement error, the smoothing basis functions are incorrectly identified, or the step jumps in the functional mean levels are neglected. To address these challenges, a factor-augmented smoothing model is proposed, and an iterative numerical estimation approach is implemented in practice. Including the factor model component in the proposed method solves the aforementioned problems since a few common factors often drive the variation that cannot be captured by the smoothing model. Asymptotic theorems are also established to demonstrate the effects of including factor structures on the smoothing results. Specifically, we show that the smoothing coefficients projected on the complement space of the factor loading matrix is asymptotically normal. As a byproduct of independent interest, an estimator for the population covariance matrix of the raw data is presented based on the proposed model. Extensive simulation studies illustrate that these factor adjustments are essential in improving estimation accuracy and avoiding the curse of dimensionality. The superiority of our model is also shown in modeling Canadian weather data and Australian temperature data.

stat.ME↗

Double bootstrapping for visualising the distribution of descriptive statistics of functional data

We propose a double bootstrap procedure for reducing coverage error in the confidence intervals of descriptive statistics for independent and identically distributed functional data. Through a series of Monte Carlo simulations, we compare the finite sample performance of single and double bootstrap procedures for estimating the distribution of descriptive statistics for independent and identically distributed functional data. At the cost of longer computational time, the double bootstrap with the same bootstrap method reduces confidence level error and provides improved coverage accuracy than the single bootstrap. Illustrated by a Canadian weather station data set, the double bootstrap procedure presents a tool for visualising the distribution of the descriptive statistics for the functional data.

stat.ME↗

Functional time series forecasting of extreme values

We consider forecasting functional time series of extreme values within a generalised extreme value distribution (GEV). The GEV distribution can be characterised using the three parameters (location, scale and shape). As a result, the forecasts of the GEV density can be accomplished by forecasting these three latent parameters. Depending on the underlying data structure, some of the three parameters can either be modelled as scalars or functions. We provide two forecasting algorithms to model and forecast these parameters. To assess the forecast uncertainty, we apply a sieve bootstrap method to construct pointwise and simultaneous prediction intervals of the forecasted extreme values. Illustrated by a daily maximum temperature dataset, we demonstrate the advantages of modelling these parameters as functions. Further, the finite-sample performance of our methods is quantified using several Monte-Carlo simulated data under a range of scenarios.

stat.ME↗

A partial least squares approach for function-on-function interaction regression

A partial least squares regression is proposed for estimating the function-on-function regression model where a functional response and multiple functional predictors consist of random curves with quadratic and interaction effects. The direct estimation of a function-on-function regression model is usually an ill-posed problem. To overcome this difficulty, in practice, the functional data that belong to the infinite-dimensional space are generally projected into a finite-dimensional space of basis functions. The function-on-function regression model is converted to a multivariate regression model of the basis expansion coefficients. In the estimation phase of the proposed method, the functional variables are approximated by a finite-dimensional basis function expansion method. We show that the partial least squares regression constructed via a functional response, multiple functional predictors, and quadratic/interaction terms of the functional predictors is equivalent to the partial least squares regression constructed using basis expansions of functional variables. From the partial least squares regression of the basis expansions of functional variables, we provide an explicit formula for the partial least squares estimate of the coefficient function of the function-on-function regression model. Because the true forms of the models are generally unspecified, we propose a forward procedure for model selection. The finite sample performance of the proposed method is examined using several Monte Carlo experiments and two empirical data analyses, and the results were found to compare favorably with an existing method.

stat.ME↗

Robust bootstrap prediction intervals for univariate and multivariate autoregressive time series models

The bootstrap procedure has emerged as a general framework to construct prediction intervals for future observations in autoregressive time series models. Such models with outlying data points are standard in real data applications, especially in the field of econometrics. These outlying data points tend to produce high forecast errors, which reduce the forecasting performances of the existing bootstrap prediction intervals calculated based on non-robust estimators. In the univariate and multivariate autoregressive time series, we propose a robust bootstrap algorithm for constructing prediction intervals and forecast regions. The proposed procedure is based on the weighted likelihood estimates and weighted residuals. Its finite sample properties are examined via a series of Monte Carlo studies and two empirical data examples.

stat.ME↗

Change point detection for COVID-19 excess deaths in Belgium

Emerging at the end of 2019, COVID-19 has become a public health threat to people worldwide. Apart from the deaths who tested positive for COVID-19, many others have died from causes indirectly related to COVID-19. Therefore, the COVID-19 confirmed deaths underestimate the influence of the pandemic on the society; instead, the measure of `excess deaths' is a more objective and comparable way to assess the scale of the epidemic and formulate lessons. One common practical issue in analyzing the impact of COVID-19 is to determine the `pre-COVID-19' period and the `post-COVID-19' period. We apply a change point detection method to identify any change points using the excess deaths in Belgium.

stat.AP↗

Bayesian bandwidth estimation for local linear fitting in nonparametric regression models

This paper presents a Bayesian sampling approach to bandwidth estimation for the local linear estimator of the regression function in a nonparametric regression model. In the Bayesian sampling approach, the error density is approximated by a location-mixture density of Gaussian densities with means the individual errors and variance a constant parameter. This mixture density has the form of a kernel density estimator of errors and is referred to as the kernel-form error density (c.f., Zhang et al., 2014). While Zhang et al. (2014) use the local constant (also known as the Nadaraya- Watson) estimator to estimate the regression function, we extend this to the local linear estimator, which produces more accurate estimation. The proposed investigation is motivated by the lack of data-driven methods for simultaneously choosing bandwidths in the local linear estimator of the regression function and kernel-form error density. Treating bandwidths as parameters, we derive an approximate (pseudo) likelihood and a posterior. A simulation study shows that the proposed bandwidth estimation outperforms the rule-of-thumb and cross-validation methods under the criterion of integrated squared errors. The proposed bandwidth estimation method is validated through a nonparametric regression model involving firm ownership concentration, and a model involving state-price density estimation.

stat.ME↗

Granger causality of bivariate stationary curve time series

We study causality between bivariate curve time series using the Granger causality generalized measures of correlation. With this measure, we can investigate which curve time series Granger-causes the other; in turn, it helps determine the predictability of any two curve time series. Illustrated by a climatology example, we find that the sea surface temperature Granger-causes the sea-level atmospheric pressure. Motivated by a portfolio management application in finance, we single out those stocks that lead or lag behind Dow-Jones industrial averages. Given a close relationship between S&P 500 index and crude oil price, we determine the leading and lagging variables.

stat.ME↗

Synergy in fertility forecasting: Improving forecast accuracy through model averaging

Accuracy in fertility forecasting has proved challenging and warrants renewed attention. One way to improve accuracy is to combine the strengths of a set of existing models through model averaging. The model-averaged forecast is derived using empirical model weights that optimise forecast accuracy at each forecast horizon based on historical data. We apply model averaging to fertility forecasting for the first time, using data for 17 countries and six models. Four model-averaging methods are compared: frequentist, Bayesian, model confidence set, and equal weights. We compute individual-model and model-averaged point and interval forecasts at horizons of one to 20 years. We demonstrate gains in average accuracy of 4-23\% for point forecasts and 3-24\% for interval forecasts, with greater gains from the frequentist and equal-weights approaches at longer horizons. Data for England \& Wales are used to illustrate model averaging in forecasting age-specific fertility to 2036. The advantages and further potential of model averaging for fertility forecasting are discussed. As the accuracy of model-averaged forecasts depends on the accuracy of the individual models, there is ongoing need to develop better models of fertility for use in forecasting and model averaging. We conclude that model averaging holds considerable promise for the improvement of fertility forecasting in a systematic way using existing models and warrants further investigation.

stat.AP↗

Forecasting Australian subnational age-specific mortality rates

When modeling sub-national mortality rates, it is important to incorporate any possible correlation among sub-populations to improve forecast accuracy. Moreover, forecasts at the sub-national level should aggregate consistently across the forecasts at the national level. In this study, we apply a grouped multivariate functional time series to forecast Australian regional and remote age-specific mortality rates and reconcile forecasts in a group structure using various methods. Our proposed method compares favorably to a grouped univariate functional time series forecasting method by comparing one-step-ahead to five-step-ahead point forecast accuracy. Thus, we demonstrate that joint modeling of sub-populations with similar mortality patterns can improve point forecast accuracy.

stat.AP↗

Retiree mortality forecasting: A partial age-range or a full age-range model?

An essential input of annuity pricing is the future retiree mortality. From observed age-specific mortality data, modeling and forecasting can be taken place in two routes. On the one hand, we can first truncate the available data to retiree ages and then produce mortality forecasts based on a partial age-range model. On the other hand, with all available data, we can first apply a full age-range model to produce forecasts and then truncate the mortality forecasts to retiree ages. We investigate the difference in modeling the logarithmic transformation of the central mortality rates between a partial age-range and a full age-range model, using data from mainly developed countries in the Human Mortality Database (2020). By evaluating and comparing the short-term point and interval forecast accuracies, we recommend the first strategy by truncating all available data to retiree ages and then produce mortality forecasts. However, when considering the long-term forecasts, it is unclear which strategy is better since it is more difficult to find a model and parameters that are optimal. This is a disadvantage of using methods based on time series extrapolation for long-term forecasting. Instead, an expectation approach, in which experts set a future target, could be considered, noting that this method has also had limited success in the past.

stat.AP↗

A comparison of Hurst exponent estimators in long-range dependent curve time series

The Hurst exponent is the simplest numerical summary of self-similar long-range dependent stochastic processes. We consider the estimation of Hurst exponent in long-range dependent curve time series. Our estimation method begins by constructing an estimate of the long-run covariance function, which we use, via dynamic functional principal component analysis, in estimating the orthonormal functions spanning the dominant sub-space of functional time series. Within the context of functional autoregressive fractionally integrated moving average models, we compare finite-sample bias, variance and mean square error among some time- and frequency-domain Hurst exponent estimators and make our recommendations.

math.ST↗

A comparison of parameter estimation in function-on-function regression

Recent technological developments have enabled us to collect complex and high-dimensional data in many scientific fields, such as population health, meteorology, econometrics, geology, and psychology. It is common to encounter such datasets collected repeatedly over a continuum. Functional data, whose sample elements are functions in the graphical forms of curves, images, and shapes, characterize these data types. Functional data analysis techniques reduce the complex structure of these data and focus on the dependences within and (possibly) between the curves. A common research question is to investigate the relationships in regression models that involve at least one functional variable. However, the performance of functional regression models depends on several factors, such as the smoothing technique, the number of basis functions, and the estimation method. This paper provides a selective comparison for function-on-function regression models where both the response and predictor(s) are functions, to determine the optimal choice of basis function from a set of model evaluation criteria. We also propose a bootstrap method to construct a confidence interval for the response function. The numerical comparisons are implemented through Monte Carlo simulations and two real data examples.

stat.ME↗

Bayesian bandwidth estimation and semi-metric selection for a functional partial linear model with unknown error density

This study examines the optimal selections of bandwidth and semi-metric for a functional partial linear model. Our proposed method begins by estimating the unknown error density using a kernel density estimator of residuals, where the regression function, consisting of parametric and nonparametric components, can be estimated by functional principal component and functional Nadayara-Watson estimators. The estimation accuracy of the regression function and error density crucially depends on the optimal estimations of bandwidth and semi-metric. A Bayesian method is utilized to simultaneously estimate the bandwidths in the regression function and kernel error density by minimizing the Kullback-Leibler divergence. For estimating the regression function and error density, a series of simulation studies demonstrate that the functional partial linear model gives improved estimation and forecast accuracies compared with the functional principal component regression and functional nonparametric regression. Using a spectroscopy dataset, the functional partial linear model yields better forecast accuracy than some commonly used functional regression models. As a by-product of the Bayesian method, a pointwise prediction interval can be obtained, and marginal likelihood can be used to select the optimal semi-metric.

stat.ME↗

Forecasting multiple functional time series in a group structure: an application to mortality

When modeling sub-national mortality rates, we should consider three features: (1) how to incorporate any possible correlation among sub-populations to potentially improve forecast accuracy through multi-population joint modeling; (2) how to reconcile sub-national mortality forecasts so that they aggregate adequately across various levels of a group structure; (3) among the forecast reconciliation methods, how to combine their forecasts to achieve improved forecast accuracy. To address these issues, we introduce an extension of grouped univariate functional time series method. We first consider a multivariate functional time series method to jointly forecast multiple related series. We then evaluate the impact and benefit of using forecast combinations among the forecast reconciliation methods. Using the Japanese regional age-specific mortality rates, we investigate one-step-ahead to 15-step-ahead point and interval forecast accuracies of our proposed extension and make recommendations.

stat.ME↗

Functional linear models for interval-valued data

Aggregation of large databases in a specific format is a frequently used process to make the data easily manageable. Interval-valued data is one of the data types that is generated by such an aggregation process. Using traditional methods to analyze interval-valued data results in loss of information, and thus, several interval-valued data models have been proposed to gather reliable information from such data types. On the other hand, recent technological developments have led to high dimensional and complex data in many application areas, which may not be analyzed by traditional techniques. Functional data analysis is one of the most commonly used techniques to analyze such complex datasets. While the functional extensions of much traditional statistical techniques are available, the functional form of the interval-valued data has not been studied well. This paper introduces the functional forms of some well-known regression models that take interval-valued data. The proposed methods are based on the function-on-function regression model, where both the response and predictor/s are functional. Through several Monte Carlo simulations and empirical data analysis, the finite sample performance of the proposed methods is evaluated and compared with the state-of-the-art.

stat.ME↗