SearcharxivSearch

arXiv subjects

Antonello Maruotti

Publications and source records attributed to Antonello Maruotti.

16 recordsLinked to original sources

Environmental Risk Assessment via Nonhomogeneous Hidden Semi-Markov Models with Penalized Vector Auto-Regression

Motivated by the study of pollution trends in the city of Bergen, we introduce a flexible statistical framework for modeling multivariate air pollution data via a nonhomogeneous Hidden Semi-Markov Vector Auto-Regression. The hidden process captures unobserved environmental conditions, while the vector autoregressive structure accounts for temporal autocorrelation and cross-pollutant dependencies. The model further allows time-varying environmental conditions to influence both the average levels of pollutant concentrations and the duration of different transient states. Parameters are estimated via maximum likelihood using a tailored Expectation-Maximization (EM) algorithm, integrated with state-specific $\ell_1$ regularization to control overfitting and automatically select relevant temporal lags. The proposal is tested on simulated data under different scenarios and then applied to daily concentrations of nitrogens and particulate matter recorded in a urban area. Environmental risk is assessed by a Shapley value-based decomposition that attribute marginal risk contributions. This approach offers a comprehensive framework for multivariate environmental risk modeling, enabling better identification of high-pollution episodes and informing policy interventions.

stat.ME

Monitoring the West-Nile virus outbreaks in Italy using open-access data

This paper introduces a comprehensive and original database on West Nile virus (WNV) outbreaks that have occurred in Italy from September 2012 to November 2022. We have digitized bulletins published by the Italian National Institute of Health (ISS) to demonstrate the potential utilization of this data for the research community. Our aim is to establish a centralized open-access repository that facilitates analysis and monitoring of the disease. We have collected and curated data on the type of infected host, along with additional information whenever available, including the type of infection, age, and geographic details at different levels of spatial aggregation. By combining our data with other sources of information such as weather data, it becomes possible to assess potential relationships between WNV outbreaks and environmental factors. We strongly believe in supporting public oversight of government epidemic management, and we emphasize that open data plays a crucial role in generating reliable results by enabling greater transparency.

stat.AP

Parsimonious Hidden Markov Models for Matrix-Variate Longitudinal Data

Hidden Markov models (HMMs) have been extensively used in the univariate and multivariate literature. However, there has been an increased interest in the analysis of matrix-variate data over the recent years. In this manuscript we introduce HMMs for matrix-variate longitudinal data, by assuming a matrix normal distribution in each hidden state. Such data are arranged in a four-way array. To address for possible overparameterization issues, we consider the spectral decomposition of the covariance matrices, leading to a total of 98 HMMs. An expectation-conditional maximization algorithm is discussed for parameter estimation. The proposed models are firstly investigated on simulated data, in terms of parameter recovery, computational times and model selection. Then, they are fitted to a four-way real data set concerning the unemployment rates of the Italian provinces, evaluated by gender and age classes, over the last 16 years.

stat.ME

Spatial modelling of COVID-19 incident cases using Richards' curve: an application to the Italian regions

We introduce an extended generalised logistic growth model for discrete outcomes, in which a network structure can be specified to deal with spatial dependence and time dependence is dealt with using an Auto-Regressive approach. A major challenge concerns the specification of the network structure, crucial to consistently estimate the canonical parameters of the generalised logistic curve, e.g. peak time and height. Parameters are estimated under the Bayesian framework, using the {\texttt{ Stan}} probabilistic programming language. The proposed approach is motivated by the analysis of the first and second wave of COVID-19 in Italy, i.e. from February 2020 to July 2020 and from July 2020 to December 2020, respectively. We analyse data at the regional level and, interestingly enough, prove that substantial spatial and temporal dependence occurred in both waves, although strong restrictive measures were implemented during the first wave. Accurate predictions are obtained, improving those of the model where independence across regions is assumed.

stat.AP

Nowcasting COVID-19 incidence indicators during the Italian first outbreak

A novel parametric regression model is proposed to fit incidence data typically collected during epidemics. The proposal is motivated by real-time monitoring and short-term forecasting of the main epidemiological indicators within the first outbreak of COVID-19 in Italy. Accurate short-term predictions, including the potential effect of exogenous or external variables are provided; this ensures to accurately predict important characteristics of the epidemic (e.g., peak time and height), allowing for a better allocation of health resources over time. Parameters estimation is carried out in a maximum likelihood framework. All computational details required to reproduce the approach and replicate the results are provided.

stat.AP

A two-part finite mixture quantile regression model for semi-continuous longitudinal data

This paper develops a two-part finite mixture quantile regression model for semi-continuous longitudinal data. The proposed methodology allows heterogeneity sources that influence the model for the binary response variable, to influence also the distribution of the positive outcomes. As is common in the quantile regression literature, estimation and inference on the model parameters are based on the Asymmetric Laplace distribution. Maximum likelihood estimates are obtained through the EM algorithm without parametric assumptions on the random effects distribution. In addition, a penalized version of the EM algorithm is presented to tackle the problem of variable selection. The proposed statistical method is applied to the well-known RAND Health Insurance Experiment dataset which gives further insights on its empirical behavior.

stat.ME

A copula-based multivariate hidden Markov model for modelling momentum in football

We investigate the potential occurrence of change points - commonly referred to as "momentum shifts" - in the dynamics of football matches. For that purpose, we model minute-by-minute in-game statistics of Bundesliga matches using hidden Markov models (HMMs). To allow for within-state correlation of the variables considered, we formulate multivariate state-dependent distributions using copulas. For the Bundesliga data considered, we find that the fitted HMMs comprise states which can be interpreted as a team showing different levels of control over a match. Our modelling framework enables inference related to causes of momentum shifts and team tactics, which is of much interest to managers, bookmakers, and sports fans.

stat.AP

An ensemble approach to short-term forecast of COVID-19 intensive care occupancy in Italian Regions

The availability of intensive care beds during the Covid-19 epidemic is crucial to guarantee the best possible treatment to severely affected patients. In this work we show a simple strategy for short-term prediction of Covid-19 ICU beds, that has proved very effective during the Italian outbreak in February to May 2020. Our approach is based on an optimal ensemble of two simple methods: a generalized linear mixed regression model which pools information over different areas, and an area-specific non-stationary integer autoregressive methodology. Optimal weights are estimated using a leave-last-out rationale. The approach has been set up and validated during the epidemic in Italy. A report of its performance for predicting ICU occupancy at Regional level is included.

stat.ME

Modelling corporate defaults: A Markov-switching Poisson log-linear autoregressive model

This article extends the autoregressive count time series model class by allowing for a model with regimes, that is, some of the parameters in the model depend on the state of an unobserved Markov chain. We develop a quasi-maximum likelihood estimator by adapting the extended Hamilton-Grey algorithm for the Poisson log-linear autoregressive model, and we perform a simulation study to check the finite sample behaviour of the estimator. The motivation for the model comes from the study of corporate defaults, in particular the study of default clustering. We provide evidence that time series of counts of US monthly corporate defaults consists of two regimes and that the so-called contagion effect, that is current defaults affect the probability of other firms defaulting in the future, is present in one of these regimes, even after controlling for financial and economic covariates. We further find evidence for that the covariate effects are different in each of the two regimes. Our results imply that the notion of contagion in the default count process is time-dependent, and thus more dynamic than previously believed.

stat.ME

On initial direction, orientation and discreteness in the analysis of circular variables

In this paper, we propose a discrete circular distribution obtained by extending the wrapped Poisson distribution. This new distribution, the Invariant Wrapped Poisson (IWP), enjoys numerous advantages: simple tractable density, parameter-parsimony and interpretability, good circular dependence structure and easy random number generation thanks to known marginal/conditional distributions. Existing discrete circular distributions strongly depend on the initial direction and orientation, i.e. a change of the reference system on the circle may lead to misleading inferential results. We investigate the invariance properties, i.e. invariance under change of initial direction and of the reference system orientation, for several continuous and discrete distributions. We prove that the introduced IWP distribution satisfies these two crucial properties. We estimate parameters in a Bayesian framework and provide all computational details to implement the algorithm. Inferential issues related to the invariance properties are discussed through numerical examples on artificial and real data.

stat.ME

Bayesian Hidden Markov Modelling Using Circular-Linear General Projected Normal Distribution

We introduce a multivariate hidden Markov model to jointly cluster time-series observations with different support, i.e. circular and linear. Relying on the general projected normal distribution, our approach allows for bimodal and/or skewed cluster-specific distributions for the circular variable. Furthermore, we relax the independence assumption between the circular and linear components observed at the same time. Such an assumption is generally used to alleviate the computational burden involved in the parameter estimation step, but it is hard to justify in empirical applications. We carry out a simulation study using different data-generation schemes to investigate model behavior, focusing on well recovering the hidden structure. Finally, the model is used to fit a real data example on a bivariate time series of wind speed and direction.

stat.AP

A flexible bivariate location-scale finite mixture approach to economic growth

We introduce a multivariate multidimensional mixed-effects regression model in a finite mixture framework. We relax the usual unidimensionality assumption on the random effects multivariate distribution. Thus, we introduce a multidimensional multivariate discrete distribution for the random terms, with a possibly different number of support points in each univariate profile, allowing for a full association structure. Our approach is motivated by the analysis of economic growth. Accordingly, we define an extended version of the augmented Solow model. Indeed, we allow all model parameters, and not only the mean, to vary according to a regression model. Moreover, we argue that countries do not follow the same growth process, and that a mixture-based approach can provide a natural framework for the detection of similar growth patterns. Our empirical findings provide evidence of heterogenous behaviors and suggest the need of a flexible approach to properly reflect the heterogeneity in the data. We further test the behavior of the proposed approach via a simulation study, considering several factors such as the number of observed units, times and levels of heterogeneity in the data.

stat.ME

Handling non-ignorable dropouts in longitudinal data: A conditional model based on a latent Markov heterogeneity structure

We illustrate a class of conditional models for the analysis of longitudinal data suffering attrition in random effects models framework, where the subject-specific random effects are assumed to be discrete and to follow a time-dependent latent process. The latent process accounts for unobserved heterogeneity and correlation between individuals in a dynamic fashion, and for dependence between the observed process and the missing data mechanism. Of particular interest is the case where the missing mechanism is non-ignorable. To deal with the topic we introduce a conditional to dropout model. A shape change in the random effects distribution is considered by directly modeling the effect of the missing data process on the evolution of the latent structure. To estimate the resulting model, we rely on the conditional maximum likelihood approach and for this aim we outline an EM algorithm. The proposal is illustrated via simulations and then applied on a dataset concerning skin cancers. Comparisons with other well-established methods are provided as well.

stat.ME

A biclustering approach to university performances: an Italian case study

University evaluation is a topic of increasing concern in Italy as well as in other countries. In empirical analysis, university activities and performances are generally measured by means of indicator variables, summarizing the available information under different perspectives. In this paper, we argue that the evaluation process is a complex issue that can not be addressed by a simple descriptive approach and thus association between indicators and similarities among the observed universities should be accounted for. Particularly, we examine faculty-level data collected from different sources, covering 55 Italian Economics faculties in the academic year 2009/2010. Making use of a clustering framework, we introduce a biclustering model that accounts for both homogeneity/heterogeneity among faculties and correlations between indicators. Our results show that there are two substantial different performances between universities which can be strictly related to the nature of the institutions, namely the Private and Public profiles . Each of the two groups has its own peculiar features and its own group-specific list of priorities, strengths and weaknesses. Thus, we suggest that caution should be used in interpreting standard university rankings as they generally do not account for the complex structure of the data.

stat.AP

Time-varying clustering of multivariate longitudinal observations

We propose a statistical method for clustering of multivariate longitudinal data into homogeneous groups. This method relies on a time-varying extension on the classical K-means algorithm, where a multivariate vector autoregressive model is additionally assumed for modeling the evolution of clusters' centroids over time. We base the inference on a least squares specification of the model and coordinate descent algorithm. To illustrate our work, we consider a longitudinal dataset on human development. Three variables are modeled, namely life expectancy, education and gross domestic product.

stat.ME

Multivariate Markov-Switching models and tail risk interdependence

Markov switching models are often used to analyze financial returns because of their ability to capture frequently observed stylized facts. In this paper we consider a multivariate Student-t version of the model as a viable alternative to the usual multivariate Gaussian distribution, providing a natural robust extension that accounts for heavy-tails and time varying non-linear correlations. Moreover, these modelling assumptions allow us to capture extreme tail co-movements which are of fundamental importance to assess the underlying dependence structure of asset returns during extreme events such as financial crisis. For the considered model we provide new risk interdependence measures which generalize the existing ones, like the Conditional Value-at-Risk (CoVaR). The proposed measures aim to capture interconnections among multiple connecting market participants which is particularly relevant during period of crisis when several institutions may contemporaneously experience distress instances. Those measures are analytically evaluated on the predictive distribution of the modes in order to provide a forward-looking risk quantification. Application on a set of U.S. banks is considered to show that the right specification of the model conditional distribution along with a multiple risk interdependence measure may help to better understand how the overall risk is shared among institutions.

stat.ME