Searcharxiv⌕ Search

arXiv subjects

Han Lin Shang

Publications and source records attributed to Han Lin Shang.

114 records · Page 7Linked to original sources

Reconciling forecasts of infant mortality rates at national and sub-national levels: Grouped time-series methods

Mortality rates are often disaggregated by different attributes, such as sex, state, education, religion or ethnicity. Forecasting mortality rates at the national and sub-national levels plays an important role in making social policies associated with the national and sub-national levels. However, base forecasts at the sub-national levels may not add up to the forecasts at the national level. To address this issue, we consider the problem of reconciling mortality rate forecasts from the viewpoint of grouped time-series forecasting methods (Hyndman et al., 2011). A bottom-up method and an optimal combination method are applied to produce point forecasts of infant mortality rates that are aggregated appropriately across the different levels of a hierarchy. We extend these two methods by considering the reconciliation of interval forecasts through a bootstrap procedure. Using the regional infant mortality rates in Australia, we investigate the one-step-ahead to 20-step-ahead point and interval forecast accuracies among the independent and these two grouped time-series forecasting methods. The proposed methods are shown to be useful for reconciling point and interval forecasts of demographic rates at the national and sub-national levels, and would be beneficial for government policy decisions regarding the allocations of current and future resources at both the national and sub-national levels.

stat.AP↗

Functional time series forecasting with dynamic updating: An application to intraday particulate matter concentration

Environmental data often take the form of a collection of curves observed sequentially over time. An example of this includes daily pollution measurement curves describing the concentration of a particulate matter in ambient air. These curves can be viewed as a time series of functions observed at equally spaced intervals over a dense grid. The nature of high-dimensional data poses challenges from a statistical aspect, due to the so-called `curse of dimensionality', but it also poses opportunities to analyze a rich source of information to better understand dynamic changes at short time intervals. Statistical methods are introduced and compared for forecasting one-day-ahead intraday concentrations of particulate matter; as new data are sequentially observed, dynamic updating methods are proposed to update point and interval forecasts to achieve better accuracy. These forecasting methods are validated through an empirical study of half-hourly concentrations of airborne particulate matter in Graz, Austria.

stat.CO↗

Mortality and life expectancy forecasting for a group of populations in developed countries: A multilevel functional data method

A multilevel functional data method is adapted for forecasting age-specific mortality for two or more populations in developed countries with high-quality vital registration systems. It uses multilevel functional principal component analysis of aggregate and population-specific data to extract the common trend and population-specific residual trend among populations. If the forecasts of population-specific residual trends do not show a long-term trend, then convergence in forecasts may be achieved. This method is first applied to age- and sex-specific data for the United Kingdom, and its forecast accuracy is then further compared with several existing methods, including independent functional data and product-ratio methods, through a multi-country comparison. The proposed method is also demonstrated by age-, sex- and state-specific data in Australia, where the convergence in forecasts can possibly be achieved by sex and state. For forecasting age-specific mortality, the multilevel functional data method is more accurate than the other coherent methods considered. For forecasting female life expectancy at birth, the multilevel functional data method is outperformed by the Bayesian method of \cite{RLG14}. For forecasting male life expectancy at birth, the multilevel functional data method performs better than the Bayesian methods in terms of point forecasts, but less well in terms of interval forecasts. Supplementary materials for this article are available online.

stat.AP↗

A plug-in bandwidth selection procedure for long run covariance estimation with stationary functional time series

In arenas of application including environmental science, economics, and medicine, it is increasingly common to consider time series of curves or functions. Many inferential procedures employed in the analysis of such data involve the long run covariance function or operator, which is analogous to the long run covariance matrix familiar to finite dimensional time series analysis and econometrics. This function may be naturally estimated using a smoothed periodogram type estimator evaluated at frequency zero that relies crucially on the choice of a bandwidth parameter. Motivated by a number of prior contributions in the finite dimensional setting, we propose a bandwidth selection method that aims to minimize the estimator's asymptotic mean squared normed error (AMSNE) in $L^2[0,1]^2$. As the AMSNE depends on unknown population quantities including the long run covariance function itself, estimates for these are plugged in in an initial step after which the estimated AMSNE can be minimized to produce an empirical optimal bandwidth. We show that the bandwidth produced in this way is asymptotically consistent with the AMSNE optimal bandwidth, with quantifiable rates, under mild stationarity and moment conditions. These results and the efficacy of the proposed methodology are evaluated by means of a comprehensive simulation study, from which we can offer practical advice on how to select the bandwidth parameter in this setting.

stat.CO↗

Selection of the optimal Box-Cox transformation parameter for modelling and forecasting age-specific fertility

The Box-Cox transformation can sometimes yield noticeable improvements in model simplicity, variance homogeneity and precision of estimation, such as in modelling and forecasting age-specific fertility. Despite its importance, there have been few studies focusing on the optimal selection of Box-Cox transformation parameters in demographic forecasting. A simple method is proposed for selecting the optimal Box-Cox transformation parameter, along with an algorithm based on an in-sample forecast error measure. Illustrated by Australian age-specific fertility, the out-of-sample accuracy of a forecasting method can be improved with the selected Box-Cox transformation parameter. Furthermore, the log transformation is not adequate for modelling and forecasting age-specific fertility. The Box-Cox transformation parameter should be embedded in statistical analysis of age-specific demographic data, in order to fully capture forecast uncertainties.

stat.AP↗

Bayesian bandwidth estimation for a nonparametric functional regression model with mixed types of regressors and unknown error density

We investigate the issue of bandwidth estimation in a nonparametric functional regression model with function-valued, continuous real-valued and discrete-valued regressors under the framework of unknown error density. Extending from the recent work of Shang (2013, Computational Statistics & Data Analysis), we approximate the unknown error density by a kernel density estimator of residuals, where the regression function is estimated by the functional Nadaraya-Watson estimator that admits mixed types of regressors. We derive a kernel likelihood and posterior density for the bandwidth parameters under the kernel-form error density, and put forward a Bayesian bandwidth estimation approach that can simultaneously estimate the bandwidths. Simulation studies demonstrated the estimation accuracy of the regression function and error density for the proposed Bayesian approach. Illustrated by a spectroscopy data set in the food quality control, we applied the proposed Bayesian approach to select the optimal bandwidths in a nonparametric functional regression model with mixed types of regressors.

stat.ME↗