SearcharxivSearch

arXiv subjects

Flavio Santi

Publications and source records attributed to Flavio Santi.

6 recordsLinked to original sources

Modeling and estimating skewed and heavy-tailed populations via unsupervised mixture models

We develop a mixture model for non-negative, heavy-tailed data, such as losses in actuarial and risk management applications. The mixture has a lognormal component, which is usually appropriate for the body of the distribution, and a Pareto-type tail, aimed at accommodating the largest observations, since the lognormal often decays too fast. Given that the tail is modeled by a zero-location Generalized Pareto distribution, the model is fully unsupervised, i.e. no threshold needs to be chosen. We show that maximum likelihood estimation can be performed by means of the EM algorithm and that the model is quite flexible in fitting data from different data-generating processes. Simulation experiments and a real-data application to automobiles claims suggest that the approach is equivalent in terms of goodness-of-fit, but easier to estimate, with respect to two existing distributions with similar features. All the methods are implemented in the R package lognGPD, available on CRAN.

stat.ME

Flexible modeling of bimodal distributions via skewed-$t$ mixtures

We propose a mixture of location-scale skewed-$t$ distributions to fit bimodal, skewed and heavy-tailed data. In particular, the mixture is based on the skewed-$t$ distribution by Fernández and Steel (1998), so that the model-building procedure can be easily extended to mixtures of other symmetric distributions. After studying the properties of the mixture, we develop a maximum likelihood estimation approach via the EM algorithm and a likelihood ratio test of the null hypothesis of no skewness in any given component. A simulation-based comparison to a recently proposed mixture of g-and-h distributions suggests that the performance of the proposed model is excellent, in terms of both estimation precision in well-specified setups and modeling capability in mis-specified frameworks. Fitting the model to the Standard & Poor's 500 distortion allows us to confirm the bimodality of its distribution, with the implication that the US stock market has historically been in bearish or bullish conditions, rather than near its fundamental value.

stat.ME

Modelling and predicting the spatio-temporal spread of Coronavirus disease 2019 (COVID-19) in Italy

Official freely available data about the number of infected at the finest possible level of spatial areal aggregation (Italian provinces) are used to model the spatio-temporal distribution of COVID-19 infections at local level. Data time horizon ranges from 26 February 20020, which is the date when the first case not directly connected with China has been discovered in northern Italy, to 18 March 2020. An endemic-epidemic multivariate time-series mixed-effects generalized linear model for areal disease counts has been implemented to understand and predict spatio-temporal diffusion of the phenomenon. Previous literature has shown that these class of models provide reliable predictions of infectious diseases in time and space. Three subcomponents characterize the estimated model. The first is related to the evolution of the disease over time; the second is characterized by transmission of the illness among inhabitants of the same province; the third remarks the effects of spatial neighbourhood and try to capture the contagion effects of nearby areas. Focusing on the aggregated time-series of the daily counts in Italy, the contribution of any of the three subcomponents do not dominate on the others and our predictions are excellent for the whole country, with an error of 3 per thousand compared to the late available data. At local level, instead, interesting distinct patterns emerge. In particular, the provinces first concerned by containment measures are those that are not affected by the effects of spatial neighbours. On the other hand, for the provinces the are currently strongly affected by contagions, the component accounting for the spatial interaction with surrounding areas is prevalent. Moreover, the proposed model provides good forecasts of the number of infections at local level while controlling for delayed reporting.

stat.AP

Reduced-bias estimation of spatial econometric models with incompletely geocoded data

The application of state-of-the-art spatial econometric models requires that the information about the spatial coordinates of statistical units is completely accurate, which is usually the case in the context of areal data. With micro-geographic point-level data, however, such information is inevitably affected by locational errors, that can be generated intentionally by the data producer for privacy protection or can be due to inaccuracy of the geocoding procedures. This unfortunate circumstance can potentially limit the use of the spatial econometric modelling framework for the analysis of micro data. Indeed, some recent contributions (see e.g. Arbia, Espa and Giuliani 2016) have shown that the presence of locational errors may have a non-negligible impact on the results. In particular, wrong spatial coordinates can lead to downward bias and increased variance in the estimation of model parameters. This contribution aims at developing a strategy to reduce the bias and produce more reliable inference for spatial econometrics models with location errors. The validity of the proposed approach is assessed by means of a Monte Carlo simulation study under different real-case scenarios. The study results show that the method is promising and can make the spatial econometric modelling of micro-geographic data possible.

stat.ME

On a property of the inequality curve $λ(p)$

The Zenga (1984) inequality curve is constant in p for Type I Pareto distributions. We show that this property holds exactly only for the Pareto distribution and, asymptotically, for distributions with power tail with index -a, with a greater than 1. Exploiting these properties one can develop powerful tools to analyze and estimate the tail of a distribution.

stat.ME

A goodness of fit test for the Pareto distribution

The Zenga (1984) inequality curve is constant in p for Type I Pareto distributions. This characterizing behavior will be exploited to obtain graphical and analytical tools for tail analysis and goodness of fit tests. A testing procedure for Pareto-type behavior based on a regression of technique will be introduced.

stat.ME