SearcharxivSearch

arXiv subjects

Han Lin Shang

Publications and source records attributed to Han Lin Shang.

At least 19 recordsLinked to original sources

Bias Correction of Long-memory Estimator of Functional Time Series via the Prefiltered Sieve Bootstrap

We investigate a bias correction procedure based on sieve bootstrapping to estimate the long-memory parameter d in stationary or nonstationary fractionally integrated processes. The resampling method implements a sieve bootstrap method on data prefiltered by a preliminary estimate of the long-memory parameter. For the initial estimate, we recommend the local polynomial Whittle with noise (LPWN) estimator in Frederiksen et al. (2012) to reduce bias, especially in the presence of a strong short-range autoregressive dependence. Through a series of simulation studies, we first highlight the issue of bias, using the local Whittle or detrended fluctuation analysis estimator. Then, we consider the LPWN estimator and show the potential improvement in bias achieved by its sieve bootstrap enhancement. As a byproduct, the sieve bootstrap can also provide confidence intervals of the memory parameter.

stat.ME

Age-specific demographic modeling and forecasting: Rolling window, expanding window, or both?

Rolling and expanding windows are widely used in age-specific demographic modeling and forecasting. Building on these approaches, we propose a simple combination method that assigns equal weight to the forecasts from both schemes. Our focus is on evaluating and comparing the forecast accuracy of the two window types in modeling age-specific mortality and fertility. Based on the multi-country comparison, the superior performance of one method often persists across different forecast horizons. In the absence of prior information, our combined approach offers a robust and practical alternative.

stat.AP

Visualizing and forecasting subnational life-table death counts: Gap forecasting methods

Subnational life-table death counts are highly correlated across time and space and differ by gender. While these associations are helpful in improving forecasts through joint modeling, less attention has been paid to identifying and understanding mortality disparities in gender and regional gaps. We propose a forecasting framework to model and forecast female life-table death counts at the national or subnational level, and to model and forecast the associated gender gap. For either females or males, we could forecast national life-table death counts and the regional gap relative to national data. By combining gender and regional gaps, we also explore the double gap by prioritizing national data and female data. Life-table death counts are unique due to their non-negativity and summability constraints. To address the constraints, we apply a one-to-one transformation, termed cumulative distribution function transformation, to obtain one- to 15-step-ahead forecasts. Using Japanese age-specific life-table death counts between ages 0 and 110+ from 1947 to 2023, we evaluate and compare point and interval forecast accuracy across gender, region, and double gaps. By focusing on these gaps, we can deepen our understanding of the possible factors driving gender and regional mortality variations.

stat.ME

Climate sensitivity analysis -- A case study from forty years of US compositional cause of death data

Emerging climate risks pose pressing challenges for insurers, governments, and businesses worldwide, as they face growing uncertainty in quantifying the impacts of climate change from physical risks. This paper aims to understand the impact of specific climate factors on mortality by cause for subgroups of the United States (US) population, through sensitivity and scenario analysis based on an increase in temperature and sea level extremes. We apply compositional data analysis (CODA) techniques to examine cause-specific deaths, treating the density of deaths as a set of dependent, non-negative values that sum to one. We couple CODA with principal component analysis on the climate factors as a means of dimension reduction, and fit generalised additive models to better reflect the non-linear relationships between the dimension-reduced principal scores and mortality by cause. The results of our analysis indicate climate-related factors have varying impacts by cause and ages within each cause, with more pronounced increases in the proportions of deaths from hypertensive heart disease as temperature and sea level extremes increase. Scenario analysis also indicates that an increase in temperature high extremes, sea level, and rainfall, in conjunction with a decrease in low temperature extremes, lead to offsetting impacts on the proportion of deaths between climate-related causes, but less offsetting by age within causes. Furthermore, the impacts on the proportions of death are more pronounced for ages between 55 and 95, reinforcing the observation that climate-related risks have a greater impact on older (and potentially more vulnerable) subgroups of the population. For life insurers specifically, these results are consistent with the natural hedge that arises between annuity and protection products, in light of increasing climate risk.

stat.AP

Conformal prediction for functional time series: Application to age-specific mortality rates

In demographic literature, forecast uncertainty is often quantified with a statistical model. This model-based approach may potentially suffer from drawbacks, namely model misspecification, selection effect, and lack of finite-sample validity. We introduce a model-agnostic and distribution-free procedure, conformal prediction, for constructing prediction intervals for a functional time series. In the family of conformal prediction, split conformal prediction divides the data into training, validation, and test sets. Within the validation set, we can select optimal tuning parameters by calibrating the empirical coverage probabilities to match their nominal values. With the selected optimal tuning parameters, we then construct the prediction intervals using the same forecasting model for the holdout data in the testing set. Without sample splitting, sequential conformal prediction sequentially updates the predicted quantiles via an autoregressive process. Using Australian age- and sex-specific log mortality rates, we evaluate and compare the interval forecast accuracy, as measured by empirical coverage probability, coverage probability difference and mean interval score, between the two variants of conformal prediction.

stat.AP

A Beta-GAM Hidden Markov Model for Proportion Time Series

We propose a hidden Markov model for univariate proportion time series taking values in (0,1), where regime switching captures latent structural changes and the emission distribution belongs to the Beta family. In each latent state, the Beta mean is linked to covariates through a generalized additive model (GAM) with spline-based smooth functions, while the Beta precision is state-specific, enabling flexible modeling of both nonlinear covariate effects and regime-dependent variability. Estimation is carried out via a penalized expectation--maximization algorithm, combining smoothing with numerical maximization of the penalized emission likelihood. To select the number of latent states and the smoothing penalty, we implement a grid search guided by standard information criteria (Akaike Information Criterion/Bayesian Information Criterion/Integrated Completed Likelihood) with a diagnostic filter that removes degenerate solutions characterized by explosive precision estimates. Uncertainty is quantified through a parametric bootstrap procedure for transition probabilities and state-dependent parameters. Simulation results demonstrate accurate recovery of transition dynamics, state precisions, and latent-state decoding. A motivating application to Russian age-specific mortality data (1960--2014, ages 0--40) illustrates how the proposed model summarizes smooth age patterns in female-to-total mortality ratios while identifying two persistent latent regimes that admit a substantive demographic interpretation in light of the country's well-documented mortality shocks that occurred over the second half of the twentieth century.

stat.ME

Robust spatial scalar-on-function regression: A Fisher-consistent redescending M-estimation approach

We develop a Fisher-consistent redescending robust estimator for the spatial scalar-on-function regression model, where a scalar response depends on both a functional predictor and a spatial autoregressive lag. Existing estimation procedures for this model are typically based on likelihood methods or monotone-loss robust M-estimators. They may be highly sensitive to vertical outliers, leverage points in the functional predictor, and numerical instability induced by strong spatial dependence. To address these issues, we propose a new estimation framework that first applies robust functional principal component analysis to obtain a contamination-resistant finite-dimensional representation of the functional predictor and then estimates the resulting spatial regression model through a bias-corrected system of M-estimating equations. The proposed method allows redescending loss functions, including Andrews' sine and Danish losses, and jointly estimates the regression coefficients, spatial dependence parameter, and scale parameter within a unified Fisher-consistent framework. For computation, we develop a hybrid IRLS-Newton algorithm that combines weighted least-squares updates for the regression parameters with a Newton-Raphson update for the spatial parameter. We establish Fisher consistency, consistency, asymptotic normality, and the asymptotic distribution of the reconstructed slope function. Monte Carlo experiments show that the proposed estimators remain competitive under clean data and substantially outperform classical and Huber-type robust competitors under contamination, particularly in severe outlier settings. An application to French air-quality data further demonstrates improved predictive performance and stable estimation of spatial dependence. Our method has been implemented in the fcsar R package.

stat.ME

Spherically Embedded Time Series with Unknown Trend and Periodic Components

Spherically embedded time series are time series with values naturally residing on or can be equivalently mapped to the sphere. Despite their ubiquity in diverse scientific fields, these data frequently exhibit complex non-stationarity driven by latent trend and periodic components. Traditional Euclidean time series methods fail to account for the intrinsic non-Euclidean geometry of the sphere, leaving a critical gap in rigorous methodologies for modelling and forecasting nonstationary spherically embedded time series. To address this methodological gap, we propose a unified geometric framework to analyse nonstationary spherically embedded time series. Central to our approach is a novel nonparametric spherical trend-periodicity decomposition model that uses an optimal-transport-based removal operation to sequentially extract the smooth trend and periodic components while preserving spherical topology. The resulting de-trended and de-seasonalised stationary residuals can be further modelled using a spherical autoregressive model, formalising a novel trend-periodic spherical autoregressive model. Theoretical foundations for the modelling procedure are established on the consistency under temporal dependence. Extensive simulations corroborate these theoretical guarantees and demonstrate the superior finite-sample predictive performance of the trend-periodic spherical autoregressive model. Finally, we validate the practical utility of our methodology through applications to electricity generation compositions and bike trip volume profiles, yielding significantly enhanced forecasting accuracy while providing interpretable insights into the underlying structural dynamics.

stat.ME

Interpretable models for forecasting high-dimensional functional time series

We study the modeling and forecasting of high-dimensional functional time series, which can be temporally dependent and cross-sectionally correlated. Central to our implementation is a functional analysis of variance by decomposing high-dimensional functional time series, such as subnational age- and sex-specific mortality observed over years, into two distinct components: a deterministic mean structure and a residual process varying over time. Unlike purely statistical dimensionality-reduction techniques, the functional analysis of variance decomposition provides an interpretable framework by partitioning the series into effects attributable to data-specific factors, such as regional and sex-level variations, and a grand functional mean. From the residual process, we implement a functional factor model to capture the remaining stochastic trends. By combining the forecasts of the residual component with the estimated deterministic structure, we obtain the forecasted curves for high-dimensional functional time series. Illustrated by the age-specific Japanese subnational mortality rates from 1975 to 2023, we evaluate and compare the accuracy of the point and interval forecasts across various forecast horizons. The results demonstrate that leveraging these interpretable components not only clarifies the underlying drivers of the data, but also improves point forecast accuracy by about 25% to 45% compared to an existing method, providing more transparent insights for evidence-based policy decisions, such as accurate modeling of financial costs of length of stay in the old-aged care facilities.

stat.ME

Attribution of Spurious Factors from High-Dimensional Functional Time Series

This article explores a general factor structure for high-dimensional nonstationary functional time series, encompassing a wide range of factor models studied in the existing literature. We investigate the asymptotic spectral behaviors of the sample covariance operator under this general data structure. A novel fundamental sufficient condition, formulated in terms of a newly introduced effective rank tailored to this setup, is established under which empirical eigen-analysis yields spurious results, rendering sample eigenvalues and eigenvectors unreliable for accurately recovering the underlying factor structure. This generalizes the results of Onatski and Wang [2021] from typical high-dimensional time series (HDTS) to the more intricate functional framework. The newly defined effective rank is rigorously analyzed through a decomposition of the effects attributable to functional factor loadings and functional factors. Contrary to the findings in the HDTS setting, empirical eigen-analysis of models with only a small number of strong non-stationary factors may still produce spurious limits in the functional framework. Therefore, additional caution is warranted when applying covariance-based statistical methods to potentially nonstationary functional data. Simulation studies are performed to determine conditions under which spurious limits occur. Real data analysis on age-specific mortality rate data from multiple locations is conducted for evidence of spurious factors induced by empirical eigen-analysis.

stat.ME

Conformal prediction for high-dimensional functional time series: Applications to subnational mortality

In statistics, forecast uncertainty is often quantified using a specified statistical model, though such approaches may be vulnerable to model misspecification, selection bias, and limited finite-sample validity. While bootstrapping can potentially mitigate some of these concerns, it is often computationally demanding. Instead, we take a model-agnostic and distribution-free approach, namely conformal prediction, to construct prediction intervals in high-dimensional functional time series. Among a rich family of conformal prediction methods, we study split and sequential conformal prediction. In split conformal prediction, the data are divided into training, validation, and test sets, where the validation set is used to select optimal tuning parameters by calibrating empirical coverage probabilities to match nominal levels; after this, prediction intervals are constructed for the test set, and their accuracy is evaluated. In contrast, sequential conformal prediction removes the need for a validation set by updating predictive quantiles sequentially via an autoregressive process. Using subnational age-specific log-mortality data from Japan and Canada, we compare the finite-sample forecast performance of these two conformal methods using empirical coverage probability and the mean interval score.

stat.ME

A Dirichlet-Multinomial-Poisson framework for the coherent analysis and forecast of cause-specific mortality

Separate modelling of cause specific mortality rates and their projections can yield inconsistent forecasts when the sum of deaths by cause does not match the total observed in a population. We develop a hierarchical probabilistic framework for cause specific mortality counts in which both the total number of deaths and the occurrence of deaths across causes are treated as random. Conditional on the total number of deaths, cause specific counts follow a multinomial distribution, whereas the total count is modelled using a Poisson distribution, and the vector of cause of death probabilities is assigned a Dirichlet distribution. The variation in cause specific mortality rates by age and calendar year is captured in both the Poisson and Dirichlet models, allowing interpretable demographic patterns while preserving coherence by construction. This model construction naturally preserves the coherence between the sum of deaths by cause and the total mortality. The method is exhibited through the analysis of cause specific mortality rates in the United States and France, sourced from the Human Mortality Database from 1979 to 2023, separately by sex and across ages, with deaths grouped into major cause categories. The empirical analysis uses a rolling 15 year out o fsample evaluation and compares the proposed model with the standard Lee Carter model and its compositional extension. The results show that coherent projections can be obtained across countries and sexes, that competitive predictive accuracy is achieved, and that uncertainty is well calibrated for both total and cause specific mortality.

stat.AP

Spherical Spatial Autoregressive Model for Spherically Embedded Spatial Data

Spherically embedded spatial data are spatially indexed observations whose values naturally reside on or can be equivalently mapped to the unit sphere. Such data are increasingly ubiquitous in fields ranging from geochemistry to demography. However, analysing such data presents unique difficulties due to the intrinsic non-Euclidean nature of the sphere, and rigorous methodologies for statistical modelling, inference, and uncertainty quantification remain limited. This paper introduces a unified framework to address these three limitations for spherically embedded spatial data. We first propose a novel spherical spatial autoregressive model that leverages optimal transport geometry and then extend it to accommodate exogenous covariates. Second, for either scenario with or without covariates, we establish the asymptotic properties of the estimators and derive a distribution-free Wald test for spatial dependence, complemented by a bootstrap procedure to enhance finite-sample performance. Third, we contribute a novel approach to uncertainty quantification by developing a conformal prediction procedure specifically tailored to spherically embedded spatial data. The practical utility of these methodological advances is illustrated through extensive simulations and applications to Spanish geochemical compositions and Japanese age-at-death mortality distributions.

stat.ME

On the Distributed Estimation for Scalar-on-Function Regression Models

This paper proposes distributed estimation procedures for three scalar-on-function regression models: the functional linear model (FLM), the functional non-parametric model (FNPM), and the functional partial linear model (FPLM). The framework addresses two key challenges in functional data analysis, namely the high computational cost of large samples and limitations on sharing raw data across institutions. Monte Carlo simulations show that the distributed estimators substantially reduce computation time while preserving high estimation and prediction accuracy for all three models. When block sizes become too small, the FPLM exhibits overfitting, leading to narrower prediction intervals and reduced empirical coverage probability. An example of an empirical study using the \textit{tecator} dataset further supports these findings.

stat.CO

Penalized spatial function-on-function regression

The function-on-function regression model is fundamental for analyzing relationships between functional covariates and responses. However, most existing function-on-function regression methodologies assume independence between observations, which is often unrealistic for spatially structured functional data. We propose a novel penalized spatial function-on-function regression model to address this limitation. Our approach extends the generalized spatial two-stage least-squares estimator to functional data, while incorporating a roughness penalty on the regression coefficient function using a tensor product of B-splines. This penalization ensures optimal smoothness, mitigating overfitting, and improving interpretability. The proposed penalized spatial two-stage least-squares estimator effectively accounts for spatial dependencies, significantly improving estimation accuracy and predictive performance. We establish the asymptotic properties of our estimator, proving its $\sqrt{n}$-consistency and asymptotic normality under mild regularity conditions. Extensive Monte Carlo simulations demonstrate the superiority of our method over existing non-penalized estimators, particularly under moderate to strong spatial dependence. In addition, an application to North Dakota weather data illustrates the practical utility of our approach in modeling spatially correlated meteorological variables. Our findings highlight the critical role of penalization in enhancing robustness and efficiency in spatial function-on-function regression models. To implement our method we used the \texttt{robflreg} package on CRAN.

stat.ME

Forecasting Australian Electricity Generation by Fuel Mix

Electricity demand and generation have become increasingly unpredictable with the growing share of variable renewable energy sources in the power system. Forecasting electricity supply by fuel mix is crucial for market operation, ensuring grid stability, optimizing costs, integrating renewable energy sources, and supporting sustainable energy planning. We introduce two statistical methods, centering on forecast reconciliation and compositional data analysis, to forecast short-term electricity supply by different types of fuel mix. Using data for five electricity markets in Australia, we study the forecast accuracy of these techniques. The bottom-up hierarchical forecasting method consistently outperforms the other approaches. Moreover, fuel mix forecasting is most accurate in power systems with a higher share of stable fossil fuel generation.

stat.AP

Weighted compositional functional data analysis for modeling and forecasting life-table death counts

Age-specific life-table death counts observed over time are examples of densities. Non-negativity and summability are constraints that sometimes require modifications of standard linear statistical methods. The centered log-ratio transformation presents a mapping from a constrained to a less constrained space. With a time series of densities, forecasts are more relevant to the recent data than the data from the distant past. We introduce a weighted compositional functional data analysis for modeling and forecasting life-table death counts. Our extension assigns higher weights to more recent data and provides a modeling scheme easily adapted for constraints. We illustrate our method using age-specific Swedish life-table death counts from 1751 to 2020. Compared to their unweighted counterparts, the weighted compositional data analytic method improves short-term point and interval forecast accuracies. The improved forecast accuracy could help actuaries improve the pricing of annuities and setting of reserves.

stat.ME

Mortality Models Ensemble via Shapley Value

Model averaging techniques in the actuarial literature aim to forecast future longevity appropriately by combining forecasts derived from various models. This approach often yields more accurate predictions than those generated by a single model. The key to enhancing forecast accuracy through model averaging lies in identifying the optimal weights from a finite sample. Utilizing sub-optimal weights in computations may adversely impact the accuracy of the model-averaged longevity forecasts. By proposing a game-theoretic approach employing Shapley values for weight selection, our study clarifies the distinct impact of each model on the collective predictive outcome. This analysis not only delineates the importance of each model in decision-making processes, but also provides insight into their contribution to the overall predictive performance of the ensemble.

stat.AP