SearcharxivSearch

arXiv subjects

Andrea Nigri

Publications and source records attributed to Andrea Nigri.

8 recordsLinked to original sources

Visualizing and forecasting subnational life-table death counts: Gap forecasting methods

Subnational life-table death counts are highly correlated across time and space and differ by gender. While these associations are helpful in improving forecasts through joint modeling, less attention has been paid to identifying and understanding mortality disparities in gender and regional gaps. We propose a forecasting framework to model and forecast female life-table death counts at the national or subnational level, and to model and forecast the associated gender gap. For either females or males, we could forecast national life-table death counts and the regional gap relative to national data. By combining gender and regional gaps, we also explore the double gap by prioritizing national data and female data. Life-table death counts are unique due to their non-negativity and summability constraints. To address the constraints, we apply a one-to-one transformation, termed cumulative distribution function transformation, to obtain one- to 15-step-ahead forecasts. Using Japanese age-specific life-table death counts between ages 0 and 110+ from 1947 to 2023, we evaluate and compare point and interval forecast accuracy across gender, region, and double gaps. By focusing on these gaps, we can deepen our understanding of the possible factors driving gender and regional mortality variations.

stat.ME

A Beta-GAM Hidden Markov Model for Proportion Time Series

We propose a hidden Markov model for univariate proportion time series taking values in (0,1), where regime switching captures latent structural changes and the emission distribution belongs to the Beta family. In each latent state, the Beta mean is linked to covariates through a generalized additive model (GAM) with spline-based smooth functions, while the Beta precision is state-specific, enabling flexible modeling of both nonlinear covariate effects and regime-dependent variability. Estimation is carried out via a penalized expectation--maximization algorithm, combining smoothing with numerical maximization of the penalized emission likelihood. To select the number of latent states and the smoothing penalty, we implement a grid search guided by standard information criteria (Akaike Information Criterion/Bayesian Information Criterion/Integrated Completed Likelihood) with a diagnostic filter that removes degenerate solutions characterized by explosive precision estimates. Uncertainty is quantified through a parametric bootstrap procedure for transition probabilities and state-dependent parameters. Simulation results demonstrate accurate recovery of transition dynamics, state precisions, and latent-state decoding. A motivating application to Russian age-specific mortality data (1960--2014, ages 0--40) illustrates how the proposed model summarizes smooth age patterns in female-to-total mortality ratios while identifying two persistent latent regimes that admit a substantive demographic interpretation in light of the country's well-documented mortality shocks that occurred over the second half of the twentieth century.

stat.ME

A Dirichlet-Multinomial-Poisson framework for the coherent analysis and forecast of cause-specific mortality

Separate modelling of cause specific mortality rates and their projections can yield inconsistent forecasts when the sum of deaths by cause does not match the total observed in a population. We develop a hierarchical probabilistic framework for cause specific mortality counts in which both the total number of deaths and the occurrence of deaths across causes are treated as random. Conditional on the total number of deaths, cause specific counts follow a multinomial distribution, whereas the total count is modelled using a Poisson distribution, and the vector of cause of death probabilities is assigned a Dirichlet distribution. The variation in cause specific mortality rates by age and calendar year is captured in both the Poisson and Dirichlet models, allowing interpretable demographic patterns while preserving coherence by construction. This model construction naturally preserves the coherence between the sum of deaths by cause and the total mortality. The method is exhibited through the analysis of cause specific mortality rates in the United States and France, sourced from the Human Mortality Database from 1979 to 2023, separately by sex and across ages, with deaths grouped into major cause categories. The empirical analysis uses a rolling 15 year out o fsample evaluation and compares the proposed model with the standard Lee Carter model and its compositional extension. The results show that coherent projections can be obtained across countries and sexes, that competitive predictive accuracy is achieved, and that uncertainty is well calibrated for both total and cause specific mortality.

stat.AP

Age-period modeling of mortality gaps: the cases of cancer and circulatory diseases

Understanding and modeling mortality patterns, especially differences in mortality rates between populations, is vital for demographic analysis and public health planning. We compare three statistical models within the age-period framework to examine differences in death counts. The models are based on the double Poisson, bivariate Poisson, and Skellam distributions, each of which provides unique strengths in capturing underlying mortality trends. Focusing on mortality data from 1960 to 2015, we analyze the two leading causes of death in Italy, which exhibit significant temporal and age-related variations. Our results reveal that the Skellam distribution offers superior accuracy and simplicity in capturing mortality differentials. These findings highlight the potential of the Skellam distribution for analyzing mortality gaps effectively.

stat.AP

Extending finite mixture models with skew-normal distributions and hidden Markov models for time series

We introduce an extension of finite mixture models by incorporating skew-normal distributions within a Hidden Markov Model framework. By assuming a constant transition probability matrix and allowing emission distributions to vary according to hidden states, the proposed model effectively captures dynamic dependencies between variables. Through the estimation of state-specific parameters, including location, scale, and skewness, the proposed model enables the detection of structural changes, such as shifts in the observed data distribution, while addressing challenges such as overfitting and computational inefficiencies inherent in Gaussian mixtures. Both simulation studies and real data analysis demonstrate the robustness and flexibility of the approach, highlighting its ability to accurately model asymmetric data and detect regime transitions. This methodological advancement broadens the applicability of a finite mixture of hidden Markov models across various fields, including demography, economics, finance, and environmental studies, offering a powerful tool for understanding complex temporal dynamics.

stat.ME

Null Distance and Temporal Functions

The notion of null distance was introduced by Sormani and Vega as part of a broader program to develop a theory of metric convergence adapted to Lorentzian geometry. Given a time function $τ$ on a spacetime $(M,g)$, the associated null distance $\hat{d}_τ$ is constructed from and closely related to the causal structure of $M$. While generally only a semi-metric, $\hat{d}_τ$ becomes a metric when $τ$ satisfies the local anti-Lipschitz condition. In this work, we focus on temporal functions, that is, differentiable functions whose gradient is everywhere past-directed timelike. Sormani and Vega showed that the class of $C^1$ temporal functions coincides with that of $C^1$ locally anti-Lipschitz time functions. When a temporal function $f$ is smooth, its level sets $M_t = f^{-1}(t)$ are spacelike hypersurfaces and thus Riemannian manifolds endowed with the induced metric $h_t$. Our main result establishes that, on any level set $M_t$ where the gradient $\nabla f$ has constant norm, the null distance $\hat{d}_f$ is bounded above by a constant multiple of the Riemannian distance $d_{h_t}$. Applying this result to a smooth regular cosmological time function $τ_g$ -- as introduced by Andersson, Galloway, and Howard -- we prove a theorem confirming a conjecture of Sakovich and Sormani (arXiv:2410.16800, 2025): if the diameters of the level sets $M_t = τ_g^{-1}(t)$ shrink to zero as $t \to 0$, then the spacetime exhibits a Big Bang singularity, as defined in their work.

math.DG

Modelling higher education dropouts using sparse and interpretable post-clustering logistic regression

Higher education dropout constitutes a critical challenge for tertiary education systems worldwide. While machine learning techniques can achieve high predictive accuracy on selected datasets, their adoption by policymakers remains limited and unsatisfactory, particularly when the objective is the unsupervised identification and characterization of student subgroups at elevated risk of dropout. The model introduced in this paper is a specialized form of logistic regression, specifically adapted to the context of university dropout analysis. Logistic regression continues to serve as a foundational tool among reliable statistical models, primarily due to the ease with which its parameters can be interpreted in terms of odds ratios. Our approach significantly extends this framework by incorporating heterogeneity within the student population. This is achieved through the application of a preliminary clustering algorithm that identifies latent subgroups, each characterized by distinct dropout propensities, which are then modeled via cluster-specific effects. We provide a detailed interpretation of the model parameters within this extended framework and enhance interpretability by imposing sparsity through a tailored variant of the LASSO algorithm. To demonstrate the practical applicability of the proposed methodology, we present an extensive case study based on the Italian university system, in which all the developed tools are systematically applied

stat.AP

Deepening Lee-Carter for longevity projections with uncertainty estimation

Undoubtedly, several countries worldwide endure to experience a continuous increase in life expectancy, extending the challenges of life actuaries and demographers in forecasting mortality. Although several stochastic mortality models have been proposed in past literature, the mortality forecasting research remains a crucial task. Recently, various research works encourage the adequacy of deep learning models to extrapolate suitable pattern within mortality data. Such a learning models allow to achieve accurate point predictions, albeit also uncertainty measures are necessary to support both model estimates reliability and risk evaluations. To the best of our knowledge, machine and deep learning literature in mortality forecasting lack for studies about uncertainty estimation. As new advance in mortality forecasting, we formalizes the deep Neural Networks integration within the Lee-Carter framework, posing a first bridge between the deep learning and the mortality density forecasts. We test our model proposal in a numerical application considering three representative countries worldwide and both genders, scrutinizing two different fitting periods. Exploiting the meaning of both biological reasonableness and plausibility of forecasts, as well as performance metrics, our findings confirm the suitability of deep learning models to improve the predictive capacity of the Lee-Carter model, providing more reliable mortality boundaries also on the long-run.

stat.AP