Searcharxiv⌕ Search

arXiv subjects

Han Lin Shang

Publications and source records attributed to Han Lin Shang.

At least 37 records · Page 2Linked to original sources

On the Distributed Estimation for Scalar-on-Function Regression Models

This paper proposes distributed estimation procedures for three scalar-on-function regression models: the functional linear model (FLM), the functional non-parametric model (FNPM), and the functional partial linear model (FPLM). The framework addresses two key challenges in functional data analysis, namely the high computational cost of large samples and limitations on sharing raw data across institutions. Monte Carlo simulations show that the distributed estimators substantially reduce computation time while preserving high estimation and prediction accuracy for all three models. When block sizes become too small, the FPLM exhibits overfitting, leading to narrower prediction intervals and reduced empirical coverage probability. An example of an empirical study using the \textit{tecator} dataset further supports these findings.

stat.CO↗

Testing for integer integration in functional time series

We develop a statistical testing procedure to examine whether the curve-valued time series of interest is integrated of order d for an integer d. The proposed procedure can distinguish between integer-integrated time series and fractionally-integrated ones, and it has broad applicability in practice. Monte Carlo simulation experiments show that the proposed testing procedure performs reasonably well. We apply our methodology to Canadian yield curve data and French sub-national age-specific mortality data. We find evidence that these time series are mostly integrated of order one, while some have fractional orders exceeding or falling below one.

stat.ME↗

AR-sieve Bootstrap for High-dimensional Time Series

This paper proposes a new AR-sieve bootstrap approach to high-dimensional time series. The major challenge of classical bootstrap methods on high-dimensional time series is two-fold: curse of dimensionality and temporal dependence. To address such a difficulty, we utilize factor modeling to reduce dimension and capture temporal dependence simultaneously. A factor-based bootstrap procedure is constructed, which performs an AR-sieve bootstrap on the extracted low-dimensional common factor time series and then recovers the bootstrap samples for the original data from the factor model. Asymptotic properties for bootstrap mean statistics and extreme eigenvalues are established. Various simulation studies further demonstrate the advantages of the new AR-sieve bootstrap in high-dimensional scenarios. An empirical application on particulate matter (PM) concentration data is studied, where bootstrap confidence intervals for mean vectors and autocovariance matrices are provided.

stat.ME↗

Interpretable additive model for analyzing high-dimensional functional time series

High-dimensional functional time series offers a powerful framework for extending functional time series analysis to settings with multiple simultaneous dimensions, capturing both temporal dynamics and cross-sectional dependencies. We propose a novel, interpretable additive model tailored for such data, designed to deliver both high predictive accuracy and clear interpretability. The model features bivariate coefficient surfaces to represent relationships across panel dimensions, with sparsity introduced via penalized smoothing and group bridge regression. This enables simultaneous estimation of the surfaces and identification of significant inter-dimensional effects. Through Monte Carlo simulations and an empirical application to Japanese subnational age-specific mortality rates, we demonstrate the proposed model's superior forecasting performance and interpretability compared to existing functional time series approaches.

stat.ME↗

Penalized spatial function-on-function regression

The function-on-function regression model is fundamental for analyzing relationships between functional covariates and responses. However, most existing function-on-function regression methodologies assume independence between observations, which is often unrealistic for spatially structured functional data. We propose a novel penalized spatial function-on-function regression model to address this limitation. Our approach extends the generalized spatial two-stage least-squares estimator to functional data, while incorporating a roughness penalty on the regression coefficient function using a tensor product of B-splines. This penalization ensures optimal smoothness, mitigating overfitting, and improving interpretability. The proposed penalized spatial two-stage least-squares estimator effectively accounts for spatial dependencies, significantly improving estimation accuracy and predictive performance. We establish the asymptotic properties of our estimator, proving its $\sqrt{n}$-consistency and asymptotic normality under mild regularity conditions. Extensive Monte Carlo simulations demonstrate the superiority of our method over existing non-penalized estimators, particularly under moderate to strong spatial dependence. In addition, an application to North Dakota weather data illustrates the practical utility of our approach in modeling spatially correlated meteorological variables. Our findings highlight the critical role of penalization in enhancing robustness and efficiency in spatial function-on-function regression models. To implement our method we used the \texttt{robflreg} package on CRAN.

stat.ME↗

Forecasting Australian Electricity Generation by Fuel Mix

Electricity demand and generation have become increasingly unpredictable with the growing share of variable renewable energy sources in the power system. Forecasting electricity supply by fuel mix is crucial for market operation, ensuring grid stability, optimizing costs, integrating renewable energy sources, and supporting sustainable energy planning. We introduce two statistical methods, centering on forecast reconciliation and compositional data analysis, to forecast short-term electricity supply by different types of fuel mix. Using data for five electricity markets in Australia, we study the forecast accuracy of these techniques. The bottom-up hierarchical forecasting method consistently outperforms the other approaches. Moreover, fuel mix forecasting is most accurate in power systems with a higher share of stable fossil fuel generation.

stat.AP↗

Weighted compositional functional data analysis for modeling and forecasting life-table death counts

Age-specific life-table death counts observed over time are examples of densities. Non-negativity and summability are constraints that sometimes require modifications of standard linear statistical methods. The centered log-ratio transformation presents a mapping from a constrained to a less constrained space. With a time series of densities, forecasts are more relevant to the recent data than the data from the distant past. We introduce a weighted compositional functional data analysis for modeling and forecasting life-table death counts. Our extension assigns higher weights to more recent data and provides a modeling scheme easily adapted for constraints. We illustrate our method using age-specific Swedish life-table death counts from 1751 to 2020. Compared to their unweighted counterparts, the weighted compositional data analytic method improves short-term point and interval forecast accuracies. The improved forecast accuracy could help actuaries improve the pricing of annuities and setting of reserves.

stat.ME↗

Mortality Models Ensemble via Shapley Value

Model averaging techniques in the actuarial literature aim to forecast future longevity appropriately by combining forecasts derived from various models. This approach often yields more accurate predictions than those generated by a single model. The key to enhancing forecast accuracy through model averaging lies in identifying the optimal weights from a finite sample. Utilizing sub-optimal weights in computations may adversely impact the accuracy of the model-averaged longevity forecasts. By proposing a game-theoretic approach employing Shapley values for weight selection, our study clarifies the distinct impact of each model on the collective predictive outcome. This analysis not only delineates the importance of each model in decision-making processes, but also provides insight into their contribution to the overall predictive performance of the ensemble.

stat.AP↗

Spatial Scalar-on-Function Quantile Regression Model

This paper introduces a novel spatial scalar-on-function quantile regression model that extends classical scalar-on-function models to account for spatial dependence and heterogeneous conditional distributions. The proposed model incorporates spatial autocorrelation through a spatially lagged response and characterizes the entire conditional distribution of a scalar outcome given a functional predictor. To address the endogeneity induced by the spatial lag term, we develop two robust estimation procedures based on instrumental variable strategies. $\sqrt{n}$-consistency and asymptotic normality of the proposed estimators are established under mild regularity conditions. We demonstrate through extensive Monte Carlo simulations that the proposed estimators outperform existing mean-based and robust alternatives, particularly in settings with strong spatial dependence and outlier contamination. We apply our method to high-resolution environmental data from the Lombardy region in Italy, using daily ozone trajectories to predict daily mean particulate matter with a diameter of less than 2.5 micrometers concentrations. The empirical results confirm the superiority of our approach in predictive accuracy, robustness, and interpretability across various quantile levels. Our method has been implemented in the \texttt{ssofqrm} R package.

stat.ME↗

A Compositional Approach to Modelling Cause-specific Mortality with Zero Counts

Understanding and forecasting mortality by cause is an essential branch of actuarial science, with wide-ranging implications for decision-makers in public policy and industry. To accurately capture trends in cause-specific mortality, it is critical to consider dependencies between causes of death and produce forecasts by age and cause coherent with aggregate mortality forecasts. One way to achieve these aims is to model cause-specific deaths using compositional data analysis (CODA), treating the density of deaths by age and cause as a set of dependent, non-negative values that sum to one. A major drawback of standard CODA methods is the challenge of zero values, which frequently occur in cause-of-death mortality modelling. Thus, we propose using a compositional power transformation, the $α$-transformation, to model cause-specific life-table death counts. The $α$-transformation offers a statistically rigorous approach to handling zero value subgroups in CODA compared to \emph{ad-hoc} techniques: adding an arbitrarily small amount. We illustrate the $α$-transformation on England and Wales, and US death counts by cause from the Human Cause-of-Death database, for cardiovascular-related causes of death. Results demonstrate the $α$-transformation improves forecast accuracy of cause-specific life-table death counts compared with log-ratio-based CODA transformations. The forecasts suggest declines in proportions of deaths from major cardiovascular causes (myocardial infarction and other ischemic heart diseases (IHD)).

stat.AP↗

Robust Functional Logistic Regression

Functional logistic regression is a popular model to capture a linear relationship between binary response and functional predictor variables. However, many methods used for parameter estimation in functional logistic regression are sensitive to outliers, which may lead to inaccurate parameter estimates and inferior classification accuracy. We propose a robust estimation procedure for functional logistic regression, in which the observations of the functional predictor are projected onto a set of finite-dimensional subspaces via robust functional principal component analysis. This dimension-reduction step reduces the outlying effects in the functional predictor. The logistic regression coefficient is estimated using an M-type estimator based on binary response and robust principal component scores. In doing so, we provide robust estimates by minimizing the effects of outliers in the binary response and functional predictor variables. Via a series of Monte-Carlo simulations and using hand radiograph data, we examine the parameter estimation and classification accuracy for the response variable. We find that the robust procedure outperforms some existing robust and non-robust methods when outliers are present, while producing competitive results when outliers are absent. In addition, the proposed method is computationally more efficient than some existing robust alternatives.

stat.ME↗

Age-period modeling of mortality gaps: the cases of cancer and circulatory diseases

Understanding and modeling mortality patterns, especially differences in mortality rates between populations, is vital for demographic analysis and public health planning. We compare three statistical models within the age-period framework to examine differences in death counts. The models are based on the double Poisson, bivariate Poisson, and Skellam distributions, each of which provides unique strengths in capturing underlying mortality trends. Focusing on mortality data from 1960 to 2015, we analyze the two leading causes of death in Italy, which exhibit significant temporal and age-related variations. Our results reveal that the Skellam distribution offers superior accuracy and simplicity in capturing mortality differentials. These findings highlight the potential of the Skellam distribution for analyzing mortality gaps effectively.

stat.AP↗

Density-valued time series: Nonparametric density-on-density regression

This paper is concerned with forecasting probability density functions. Density functions are nonnegative and have a constrained integral; thus, they do not constitute a vector space. Implementing unconstrained functional time-series forecasting methods is problematic for such nonlinear and constrained data. A novel forecasting method is developed based on a nonparametric function-on-function regression, where both the response and the predictor are probability density functions. Asymptotic properties of our nonparametric regression estimator are established, as well as its finite-sample performance through a series of Monte-Carlo simulation studies. Using COVID-19 data from the French department and age-specific period life tables from the United States, we assess and compare the finite-sample forecast accuracy of the proposed method with several existing methods.

stat.ME↗

On function-on-function linear quantile regression

We present two innovative functional partial quantile regression algorithms designed to accurately and efficiently estimate the regression coefficient function within the function-on-function linear quantile regression model. Our algorithms utilize functional partial quantile regression decomposition to effectively project the infinite-dimensional response and predictor variables onto a finite-dimensional space. Within this framework, the partial quantile regression components are approximated using a basis expansion approach. Consequently, we approximate the infinite-dimensional function-on-function linear quantile regression model using a multivariate quantile regression model constructed from these partial quantile regression components. To evaluate the efficacy of our proposed techniques, we conduct a series of Monte Carlo experiments and analyze an empirical dataset, demonstrating superior performance compared to existing methods in finite-sample scenarios. Our techniques have been implemented in the ffpqr package in R.

stat.ME↗

Forecasting intraday particle number size distribution: A functional time series approach

Particulate matter data now include various particle sizes, which often manifest as a collection of curves observed sequentially over time. When considering 51 distinct particle sizes, these curves form a high-dimensional functional time series observed over equally spaced and densely sampled grids. While high dimensionality poses statistical challenges due to the curse of dimensionality, it also offers a rich source of information that enables detailed analysis of temporal variation across short time intervals for all particle sizes. To model this complexity, we propose a multilevel functional time series framework incorporating a functional factor model to facilitate one-day-ahead forecasting. To quantify forecast uncertainty, we develop a calibration approach and a split conformal prediction approach to construct prediction intervals. Both approaches are designed to minimise the absolute difference between empirical and nominal coverage probabilities using a validation dataset. Furthermore, to improve forecast accuracy as new intraday data become available, we implement dynamic updating techniques for point and interval forecasts. The proposed methods are validated through an empirical application to hourly measurements of particulate matter in 51 size categories in London.

stat.ME↗

Intraday FX Volatility-Curve Forecasting with Functional GARCH Approaches

This paper seeks to forecast intraday volatility curves for major foreign exchange (FX) currencies using functional GARCH models. Intraday return curves are observed at a daily frequency, yet preserve the full high-frequency trading structure, enabling volatility analysis at the intraday level. We demonstrate that the USD/EUR, USD/GBP, and USD/JPY intraday return curves exhibit strong cross-dependence, while individually they are serially uncorrelated but display long-range conditional heteroskedasticity. Embedding cross-currency dependence via multi-level functional principal component analysis and adding intraday bid-ask spread curves as exogenous drivers significantly improves intraday and day-ahead volatility forecasts relative to functional and realised-volatility baselines. The precise volatility forecasts motivate the construction of intraday Value-at-Risk (VaR). An intraday risk management application highlights that predicted intraday VaR curves can help mitigate dramatic losses in intraday trading strategies, showcasing their practical economic benefits in FX markets.

stat.ME↗

Extending finite mixture models with skew-normal distributions and hidden Markov models for time series

We introduce an extension of finite mixture models by incorporating skew-normal distributions within a Hidden Markov Model framework. By assuming a constant transition probability matrix and allowing emission distributions to vary according to hidden states, the proposed model effectively captures dynamic dependencies between variables. Through the estimation of state-specific parameters, including location, scale, and skewness, the proposed model enables the detection of structural changes, such as shifts in the observed data distribution, while addressing challenges such as overfitting and computational inefficiencies inherent in Gaussian mixtures. Both simulation studies and real data analysis demonstrate the robustness and flexibility of the approach, highlighting its ability to accurately model asymmetric data and detect regime transitions. This methodological advancement broadens the applicability of a finite mixture of hidden Markov models across various fields, including demography, economics, finance, and environmental studies, offering a powerful tool for understanding complex temporal dynamics.

stat.ME↗

Constructing prediction intervals for the age distribution of deaths

We introduce a model-agnostic procedure to construct prediction intervals for the age distribution of deaths. The age distribution of deaths is an example of constrained data, which are nonnegative and have a constrained integral. A centered log-ratio transformation and a cumulative distribution function transformation are used to remove the two constraints, where the latter transformation can also handle the presence of zero counts. Our general procedure divides data samples into training, validation, and testing sets. Within the validation set, we can select an optimal tuning parameter by calibrating the empirical coverage probabilities to be close to their nominal ones. With the selected optimal tuning parameter, we then construct the pointwise prediction intervals using the same models for the holdout data in the testing set. Using Japanese age- and sex-specific life-table death counts, we assess and evaluate the interval forecast accuracy with a suite of functional time-series models.

stat.ME↗