SearcharxivSearch

arXiv subjects

Carlo Gaetan

Publications and source records attributed to Carlo Gaetan.

13 recordsLinked to original sources

Continuous mixtures of Gaussian processes as models for spatial extremes

Spatial modelling of extreme values allows studying the risk of joint occurrence of extreme events at different locations and is of significant interest in climatic and other environmental sciences. A popular class of dependence models for spatial extremes is that of random location-scale mixtures, in which a spatial "baseline" process is multiplied or shifted by a random variable, potentially altering its extremal dependence behaviour. Gaussian location-scale mixtures retain benefits of their Gaussian baseline processes while overcoming some of their limitations, such as symmetry, light tails and weak tail dependence. We review properties of Gaussian location-scale mixtures and develop novel constructions with interesting features, together with a general algorithm for conditional simulation from these models. We leverage their flexibility to propose extended extreme-value models, that allow for appropriately modelling not only the tails but also the bulk of the data. This is important in many applications and avoids the need to explicitly select the events considered as extreme. We propose new solutions for likelihood inference in parametric models of Gaussian location-scale mixtures, in order to avoid the numerical bottleneck given by the latent location and scale variables that can lead to high computational cost of standard likelihood evaluations. The effectiveness of the models and of the inference methods is confirmed with simulated data examples, and we present an application to wildfire-related weather variables in Portugal. Although not detailed here, the approaches would also be straightforward to use for modelling multivariate (non spatial) data.

stat.ME

A parsimonious tail compliant multiscale statistical model for aggregated rainfall

Modeling rainfall intensity distributions across aggregation scales (from sub-hourly to weekly) is essential for hydrological risk analysis and IDF curves. Aggregation naturally imposes mathematical constraints: return levels must be ordered by time scale, as daily accumulations necessarily exceed sub-daily ones. From a statistical perspective, each aggregation step should ideally not require additional parameters, yet parsimonious models describing the full distribution remain scarce, as most literature focuses on seasonal block maxima. In this study, we propose a parsimonious framework to model all rainfall intensities (low to large) across scales. We utilize the Extended Generalized Pareto Distribution (EGPD), which aligns with extreme value theory for both tails while remaining flexible for the bulk of the distribution. We establish a general result on the behavior of EGPD variables under various aggregation procedures. To overcome the difficulty of direct likelihood inference, we link the EGPD class to Poisson compound sums. This allows the use of the Panjer algorithm for efficient composite likelihood evaluation. Our approach ensures that return levels do not cross across scales and enables estimation for return periods below annual or seasonal levels. We demonstrate the method using sub-hourly series from six French stations with diverse climates. Only eight parameters are needed per station to capture scales from six minutes to three days. IDF curves above and below the annual scale are provided.

stat.AP

Multivariate distributional modeling of low, moderate, and large intensities without threshold selection steps

In fields such as hydrology and climatology, modelling the entire distribution of positive data is essential, as stakeholders require insights into the full range of values, from low to extreme. Traditional approaches often segment the distribution into separate regions, which introduces subjectivity and limits coherence. This is especially true when dealing with multivariate data. In line with multivariate extreme value theory, this paper presents a unified, threshold-free framework for modelling marginal behaviours and dependence structures based on an extended generalized Pareto distribution (EGPD). We propose decomposing multivariate data into radial and angular components. The radial component is modelled using a semi-parametric EGPD and the angular distribution is permitted to vary conditionally. This approach allows for sufficiently flexible dependence modelling. The hierarchical structure of the model facilitates the inference process. First, we combine classical maximum likelihood estimation (MLE) methods with semi-parametric approaches based on Bernstein polynomials to estimate the distribution of the radial component. Then, we use multivariate regression techniques to estimate the angular component's parameters. The model is evaluated through synthetic simulations and applied to hydrological datasets to exemplify its capacity to capture heavy-tailed marginals and complex multivariate dependencies without threshold specification.

stat.ME

Joint modeling of low and high extremes using a multivariate extended generalized Pareto distribution

In most risk assessment studies, it is important to accurately capture the entire distribution of the multivariate random vector of interest from low to high values. For example, in climate sciences, low precipitation events may lead to droughts, while heavy rainfall may generate large floods, and both of these extreme scenarios can have major impacts on the safety of people and infrastructure, as well as agricultural or other economic sectors. In the univariate case, the extended generalized Pareto distribution (eGPD) was specifically developed to accurately model low, moderate, and high precipitation intensities, while bypassing the threshold selection procedure usually conducted in extreme-value analyses. In this work, we extend this approach to the multivariate case. The proposed multivariate eGPD has the following appealing properties: (1) its marginal distributions behave like univariate eGPDs; (2) its lower and upper joint tails comply with multivariate extreme-value theory, with key parameters separately controlling dependence in each joint tail; and (3) the model allows for fast simulation and is thus amenable to simulation-based inference. We propose estimating model parameters by leveraging modern neural approaches, where a neural network, once trained, can provide point estimates, credible intervals, or full posterior approximations in a fraction of a second. Our new methodology is illustrated by application to daily rainfall times series data from the Netherlands. The proposed model is shown to provide satisfactory marginal and dependence fits from low to high quantiles.

stat.ME

Flexible space-time models for extreme data

Extreme value analysis is an essential methodology in the study of rare and extreme events, which hold significant interest in various fields, particularly in the context of environmental sciences. Models that employ the exceedances of values above suitably selected high thresholds possess the advantage of capturing the "sub-asymptotic" dependence of data. This paper presents an extension of spatial random scale mixture models to the spatio-temporal domain. A comprehensive framework for characterizing the dependence structure of extreme events across both dimensions is provided. Indeed, the model is capable of distinguishing between asymptotic dependence and independence, both in space and time, through the use of parametric inference. The high complexity of the likelihood function for the proposed model necessitates a simulation approach based on neural networks for parameter estimation, which leverages summaries of the sub-asymptotic dependence present in the data. The effectiveness of the model in assessing the limiting dependence structure of spatio-temporal processes is demonstrated through both simulation studies and an application to rainfall datasets.

stat.ME

An extended generalized Pareto regression model for count data

The statistical modeling of discrete extremes has received less attention than their continuous counterparts in the Extreme Value Theory (EVT) literature. One approach to the transition from continuous to discrete extremes is the modeling of threshold exceedances of integer random variables by the discrete version of the generalized Pareto distribution. However, the optimal choice of thresholds defining exceedances remains a problematic issue. Moreover, in a regression framework, the treatment of the majority of non-extreme data below the selected threshold is either ignored or separated from the extremes. To tackle these issues, we expand on the concept of employing a smooth transition between the bulk and the upper tail of the distribution. In the case of zero inflation, we also develop models with an additional parameter. To incorporate possible predictors, we relate the parameters to additive smoothed predictors via an appropriate link, as in the generalized additive model (GAM) framework. A penalized maximum likelihood estimation procedure is implemented. We illustrate our modeling proposal with a real dataset of avalanche activity in the French Alps. With the advantage of bypassing the threshold selection step, our results indicate that the proposed models are more flexible and robust than competing models, such as the negative binomial distribution

stat.ME

Spatial quantile clustering of climate data

In the era of climate change, the distribution of climate variables evolves with changes not limited to the mean value. Consequently, clustering algorithms based on central tendency could produce misleading results when used to summarize spatial and/or temporal patterns. We present a novel approach to spatial clustering of time series based on quantiles using a Bayesian framework that incorporates a spatial dependence layer based on a Markov random field. A series of simulations tested the proposal, then applied to the sea surface temperature of the Mediterranean Sea, one of the first seas to be affected by the effects of climate change.

stat.ME

Distributional regression models for Extended Generalized Pareto distributions

The Extended Generalized Pareto Distribution (EGPD) (Naveau et al. 2016) is a family of distribution that has been introduced to model the full range of a positive random variable but with the lower and the upper tails distributed according to the peaks-over-threshold methodology. The aim of this article is to augment the scope of application of EGPD allowing the analyst to incorporate the effect of covariates on the model. In particular we introduce a specification where the parameters of EGPD can be modeled as additive functions of the covariates, e.g. space or time. As a related product we provide an add-on code written in R that it is flexible enough to implement the EGPD in a generic way, allowing to introduce new parametric forms. We show the potential of our add-on on the modeling of hourly rainfalls over the North-West region of France and discuss modeling strategies.

stat.ME

Clustering of bivariate satellite time series: a quantile approach

Clustering has received much attention in Statistics and Machine learning with the aim of developing statistical models and autonomous algorithms which are capable of acquiring information from raw data in order to perform exploratory analysis.Several techniques have been developed to cluster sampled univariate vectors only considering the average value over the whole period and as such they have not been able to explore fully the underlying distribution as well as other features of the data, especially in presence of structured time series. We propose a model-based clustering technique that is based on quantile regression permitting us to cluster bivariate time series at different quantile levels. We model the within cluster density using asymmetric Laplace distribution allowing us to take into account asymmetry in the distribution of the data. We evaluate the performance of the proposed technique through a simulation study. The method is then applied to cluster time series observed from Glob-colour satellite data related to trophic status indices with aim of evaluating their temporal dynamics in order to identify homogeneous areas, in terms of trophic status, in the Gulf of Gabes.

stat.ME

Modeling and simulating depositional sequences using latent Gaussian random fields

Simulating a depositional (or stratigraphic) sequence conditionally on borehole data is a long-standing problem in hydrogeology and in petroleum geostatistics. This paper presents a new rule-based approach for simulating depositional sequences of surfaces conditionally on lithofacies thickness data. The thickness of each layer is modeled by a transformed latent Gaussian random field allowing for null thickness thanks to a truncation process. Layers are sequentially stacked above each other following the regional stratigraphic sequence. By choosing adequately the variograms of these random fields, the simulated surfaces separating two layers can be continuous and smooth. Borehole information is often incomplete in the sense that it does not provide direct information as to the exact layer some observed thickness belongs to. The latent Gaussian model proposed in this paper offers a natural solution to this problem by means of a Bayesian setting with a Markov Chain Monte Carlo (MCMC) algorithm that can explore all possible configurations compatible with the data. The model and the associated MCMC algorithm are validated on synthetic data and then applied to a subsoil in the Venetian Plain with a moderately dense network of cored boreholes.

stat.AP

Hierarchical space-time modeling of exceedances with an application to rainfall data

The statistical modeling of space-time extremes in environmental applications is key to understanding complex dependence structures in original event data and to generating realistic scenarios for impact models. In this context of high-dimensional data, we propose a novel hierarchical model for high threshold exceedances defined over continuous space and time by embedding a space-time Gamma process convolution for the rate of an exponential variable, leading to asymptotic independence in space and time. Its physically motivated anisotropic dependence structure is based on geometric objects moving through space-time according to a velocity vector. We demonstrate that inference based on weighted pairwise likelihood is fast and accurate. The usefulness of our model is illustrated by an application to hourly precipitation data from a study region in Southern France, where it clearly improves on an alternative censored Gaussian space-time random field model. While classical limit models based on threshold-stability fail to appropriately capture relatively fast joint tail decay rates between asymptotic dependence and classical independence, strong empirical evidence from our application and other recent case studies motivates the use of more realistic asymptotic independence models such as ours.

stat.ME

Comparing composite likelihood methods based on pairs for spatial Gaussian random fieldsM

In the last years there has been a growing interest in proposing methods for estimating covariance functions for geostatistical data. Among these, maximum likelihood estimators have nice features when we deal with a Gaussian model. However maximum likelihood becomes impractical when the number of observations is very large. In this work we review some solutions and we contrast them in terms of loss of statistical efficiency and computational burden. Specifically we focus on three types of weighted composite likelihood functions based on pairs and we compare them with the method of covariance tapering. Asymptotics properties of the three estimation methods are derived. We illustrate the effectiveness of the methods through theoretical examples, simulation experiments and by analysing a data set on yearly total precipitation anomalies at weather stations in the United States.

stat.ME

Estimation of spatial max-stable models using threshold exceedances

Parametric inference for spatial max-stable processes is difficult since the related likelihoods are unavailable. A composite likelihood approach based on the bivariate distribution of block maxima has been recently proposed in the literature. However modeling block maxima is a wasteful approach provided that other information is available. Moreover an approach based on block, typically annual, maxima is unable to take into account the fact that maxima occur or not simultaneously. If time series of, say, daily data are available, then estimation procedures based on exceedances of a high threshold could mitigate such problems. In this paper we focus on two approaches for composing likelihoods based on pairs of exceedances. The first one comes from the tail approximation for bivariate distribution proposed by Ledford and Tawn (1996) when both pairs of observations exceed the fixed threshold. The second one uses the bivariate extension (Rootzen and Tajvidi, 2006) of the generalized Pareto distribution which allows to model exceedances when at least one of the components is over the threshold. The two approaches are compared through a simulation study according to different degrees of spatial dependency. Results show that both the strength of the spatial dependencies and the threshold choice play a fundamental role in determining which is the best estimating procedure.

stat.AP