SearcharxivSearch

arXiv subjects

Thomas Opitz

Publications and source records attributed to Thomas Opitz.

At least 37 records · Page 2Linked to original sources

Modeling spatial extremes using normal mean-variance mixtures

Classical models for multivariate or spatial extremes are mainly based upon the asymptotically justified max-stable or generalized Pareto processes. These models are suitable when asymptotic dependence is present, i.e., the joint tail decays at the same rate as the marginal tail. However, recent environmental data applications suggest that asymptotic independence is equally important and, unfortunately, existing spatial models in this setting that are both flexible and can be fitted efficiently are scarce. Here, we propose a new spatial copula model based on the generalized hyperbolic distribution, which is a specific normal mean-variance mixture and is very popular in financial modeling. The tail properties of this distribution have been studied in the literature, but with contradictory results. It turns out that the proofs from the literature contain mistakes. We here give a corrected theoretical description of its tail dependence structure and then exploit the model to analyze a simulated dataset from the inverted Brown-Resnick process, hindcast significant wave height data in the North Sea, and wind gust data in the state of Oklahoma, USA. We demonstrate that our proposed model is flexible enough to capture the dependence structure not only in the tail but also in the bulk.

stat.ME

Spatial hierarchical modeling of threshold exceedances using rate mixtures

We develop new flexible univariate models for light-tailed and heavy-tailed data, which extend a hierarchical representation of the generalized Pareto (GP) limit for threshold exceedances. These models can accommodate departure from asymptotic threshold stability in finite samples while keeping the asymptotic GP distribution as a special (or boundary) case and can capture the tails and the bulk jointly without losing much flexibility. Spatial dependence is modeled through a latent process, while the data are assumed to be conditionally independent. Focusing on a gamma-gamma model construction, we design penalized complexity priors for crucial model parameters, shrinking our proposed spatial Bayesian hierarchical model toward a simpler reference whose marginal distributions are GP with moderately heavy tails. Our model can be fitted in fairly high dimensions using Markov chain Monte Carlo by exploiting the Metropolis-adjusted Langevin algorithm (MALA), which guarantees fast convergence of Markov chains with efficient block proposals for the latent variables. We also develop an adaptive scheme to calibrate the MALA tuning parameters. Moreover, our model avoids the expensive numerical evaluations of multifold integrals in censored likelihood expressions. We demonstrate our new methodology by simulation and application to a dataset of extreme rainfall events that occurred in Germany. Our fitted gamma-gamma model provides a satisfactory performance and can be successfully used to predict rainfall extremes at unobserved locations.

stat.ME

Modeling Non-Stationary Temperature Maxima Based on Extremal Dependence Changing with Event Magnitude

The modeling of spatio-temporal trends in temperature extremes can help better understand the structure and frequency of heatwaves in a changing climate. Here, we study annual temperature maxima over Southern Europe using a century-spanning dataset observed at 44 monitoring stations. Extending the spectral representation of max-stable processes, our modeling framework relies on a novel construction of max-infinitely divisible processes, which include covariates to capture spatio-temporal non-stationarities. Our new model keeps a popular max-stable process on the boundary of the parameter space, while flexibly capturing weakening extremal dependence at increasing quantile levels and asymptotic independence. This is achieved by linking the overall magnitude of a spatial event to its spatial correlation range, in such a way that more extreme events become less spatially dependent, thus more localized. Our model reveals salient features of the spatio-temporal variability of European temperature extremes, and it clearly outperforms natural alternative models. Results show that the spatial extent of heatwaves is smaller for more severe events at higher altitudes, and that recent heatwaves are moderately wider. Our probabilistic assessment of the 2019 annual maxima confirms the severity of the 2019 heatwaves both spatially and at individual sites, especially when compared to climatic conditions prevailing in 1950-1975.

stat.ME

High-resolution Bayesian mapping of landslide hazard with unobserved trigger event

Statistical models for landslide hazard enable mapping of risk factors and landslide occurrence intensity by using geomorphological covariates available at high spatial resolution. However, the spatial distribution of the triggering event (e.g., precipitation or earthquakes) is often not directly observed. In this paper, we develop Bayesian spatial hierarchical models for point patterns of landslide occurrences using different types of log-Gaussian Cox processes. Starting from a competitive baseline model that captures the unobserved precipitation trigger through a spatial random effect at slope unit resolution, we explore novel complex model structures that take clusters of events arising at small spatial scales into account, as well as nonlinear or spatially-varying covariate effects. For a 2009 event of around 4000 precipitation-triggered landslides in Sicily, Italy, we show how to fit our proposed models efficiently using the integrated nested Laplace approximation (INLA), and rigorously compare the performance of our models both from a statistical and applied perspective. In this context, we argue that model comparison should not be based on a single criterion, and that different models of various complexity may provide insights into complementary aspects of the same applied problem. In our application, our models are found to have mostly the same spatial predictive performance, implying that key to successful prediction is the inclusion of a slope-unit resolved random effect capturing the precipitation trigger. Interestingly, a parsimonious formulation of space-varying slope effects reflects a physical interpretation of the precipitation trigger: in subareas with weak trigger, the slope steepness is shown to be mostly irrelevant.

stat.ME

Max-infinitely divisible models and inference for spatial extremes

For many environmental processes, recent studies have shown that the dependence strength is decreasing when quantile levels increase. This implies that the popular max-stable models are inadequate to capture the rate of joint tail decay, and to estimate joint extremal probabilities beyond observed levels. We here develop a more flexible modeling framework based on the class of max-infinitely divisible processes, which extend max-stable processes while retaining dependence properties that are natural for maxima. We propose two parametric constructions for max-infinitely divisible models, which relax the max-stability property but remain close to some popular max-stable models obtained as special cases. The first model considers maxima over a finite, random number of independent observations, while the second model generalizes the spectral representation of max-stable processes. Inference is performed using a pairwise likelihood. We illustrate the benefits of our new modeling framework on Dutch wind gust maxima calculated over different time units. Results strongly suggest that our proposed models outperform other natural models, such as the Student-t copula process and its max-stable limit, even for large block sizes.

stat.ME

Semi-parametric resampling with extremes

Nonparametric resampling methods such as Direct Sampling are powerful tools to simulate new datasets preserving important data features such as spatial patterns from observed datasets while using only minimal assumptions. However, such methods cannot generate extreme events beyond the observed range of data values. We here propose using tools from extreme value theory for stochastic processes to extrapolate observed data towards yet unobserved high quantiles. Original data are first enriched with new values in the tail region, and then classical resampling algorithms are applied to enriched data. In a first approach to enrichment that we label "naive resampling", we generate an independent sample of the marginal distribution while keeping the rank order of the observed data. We point out inaccuracies of this approach around the most extreme values, and therefore develop a second approach that works for datasets with many replicates. It is based on the asymptotic representation of extreme events through two stochastically independent components: a magnitude variable, and a profile field describing spatial variation. To generate enriched data, we fix a target range of return levels of the magnitude variable, and we resample magnitudes constrained to this range. We then use the second approach to generate heatwave scenarios of yet unobserved magnitude over France, based on daily temperature reanalysis training data for the years 2010 to 2016.

stat.ME

Bayesian space-time gap filling for inference on extreme hot-spots: an application to Red Sea surface temperatures

We develop a method for probabilistic prediction of extreme value hot-spots in a spatio-temporal framework, tailored to big datasets containing important gaps. In this setting, direct calculation of summaries from data, such as the minimum over a space-time domain, is not possible. To obtain predictive distributions for such cluster summaries, we propose a two-step approach. We first model marginal distributions with a focus on accurate modeling of the right tail and then, after transforming the data to a standard Gaussian scale, we estimate a Gaussian space-time dependence model defined locally in the time domain for the space-time subregions where we want to predict. In the first step, we detrend the mean and standard deviation of the data and fit a spatially resolved generalized Pareto distribution to apply a correction of the upper tail. To ensure spatial smoothness of the estimated trends, we either pool data using nearest-neighbor techniques, or apply generalized additive regression modeling. To cope with high space-time resolution of data, the local Gaussian models use a Markov representation of the Matérn correlation function based on the stochastic partial differential equations (SPDE) approach. In the second step, they are fitted in a Bayesian framework through the integrated nested Laplace approximation implemented in R-INLA. Finally, posterior samples are generated to provide statistical inferences through Monte-Carlo estimation. Motivated by the 2019 Extreme Value Analysis data challenge, we illustrate our approach to predict the distribution of local space-time minima in anomalies of Red Sea surface temperatures, using a gridded dataset (11315 days, 16703 pixels) with artificially generated gaps. In particular, we show the improved performance of our two-step approach over a purely Gaussian model without tail transformations.

stat.ME

Landscape allocation: stochastic generators and statistical inference

In agricultural landscapes, the composition and spatial configuration of cultivated and semi-natural elements strongly impact species dynamics, their interactions and habitat connectivity. To allow for landscape structural analysis and scenario generation, we here develop statistical tools for real landscapes composed of geometric elements including 2D patches but also 1D linear elements such as hedges. We design generative stochastic models that combine a multiplex network representation and Gibbs energy terms to characterize the distributional behavior of landscape descriptors for land-use categories. We implement Metropolis-Hastings for this new class of models to sample agricultural scenarios featuring parameter-controlled spatial and temporal patterns (e.g., geometry, connectivity, crop-rotation). Pseudolikelihood-based inference allows studying the relevance of model components in real landscapes through statistical and functional validation, the latter achieved by comparing commonly used landscape metrics between observed and simulated landscapes. Models fitted to subregions of the Lower Durance Valley (France) indicate strong deviation from random allocation, and they realistically capture small-scale landscape patterns. In summary, our approach of statistical modeling improves the understanding of structural and functional aspects of agro-ecosystems, and it enables simulation-based theoretical analysis of how landscape patterns shape biological and ecological processes.

q-bio.PE

Point-process based Bayesian modeling of space-time structures of forest fire occurrences in Mediterranean France

Due to climate change and human activity, wildfires are expected to become more frequent and extreme worldwide, causing economic and ecological disasters. The deployment of preventive measures and operational forecasts can be aided by stochastic modeling that helps to understand and quantify the mechanisms governing the occurrence intensity. We here develop a point process framework for wildfire ignition points observed in the French Mediterranean basin since 1995, and we fit a spatio-temporal log-Gaussian Cox process with monthly temporal resolution in a Bayesian framework using the integrated nested Laplace approximation (INLA). Human activity is the main direct cause of wildfires and is indirectly measured through a number of appropriately defined proxies related to land-use covariates (urbanization, road network) in our approach, and we further integrate covariates of climatic and environmental conditions to explain wildfire occurrences. We include spatial random effects with Matérn covariance and temporal autoregression at yearly resolution. Two major methodological challenges are tackled: first, handling and unifying multi-scale structures in data is achieved through computer-intensive preprocessing steps with GIS software and kriging techniques; second, INLA-based estimation with high-dimensional response vectors and latent models is facilitated through intra-year subsampling, taking into account the occurrence structure of wildfires.

stat.AP

Space-Time Landslide Predictive Modelling

Landslides are nearly ubiquitous phenomena and pose severe threats to people, properties, and the environment. Investigators have for long attempted to estimate landslide hazard to determine where, when, and how destructive landslides are expected to be in an area. This information is useful to design landslide mitigation strategies, and to reduce landslide risk and societal and economic losses. In the geomorphology literature, most attempts at predicting the occurrence of populations of landslides rely on the observation that landslides are the result of multiple interacting, conditioning and triggering factors. Here, we propose a novel Bayesian modelling framework for the prediction of space-time landslide occurrences of the slide type caused by weather triggers. We consider log-Gaussian cox processes, assuming that individual landslides stem from a point process described by an unknown intensity function. We tested our prediction framework in the Collazzone area, Umbria, Central Italy, for which a detailed multi-temporal landslide inventory spanning 1941-2014 is available together with lithological and bedding data. We tested five models of increasing complexity. Our most complex model includes fixed effects and latent spatio-temporal effects, thus largely fulfilling the common definition of landslide hazard in the literature. We quantified the spatio-temporal predictive skill of our model and found that it performed better than simpler alternatives. We then developed a novel classification strategy and prepared an intensity-susceptibility landslide map, providing more information than traditional susceptibility zonations for land planning and management. We expect our novel approach to lead to better projections of future landslides, and to improve our collective understanding of the evolution of landscapes dominated by mass-wasting processes under geophysical and weather triggers.

stat.AP

Hierarchical space-time modeling of exceedances with an application to rainfall data

The statistical modeling of space-time extremes in environmental applications is key to understanding complex dependence structures in original event data and to generating realistic scenarios for impact models. In this context of high-dimensional data, we propose a novel hierarchical model for high threshold exceedances defined over continuous space and time by embedding a space-time Gamma process convolution for the rate of an exponential variable, leading to asymptotic independence in space and time. Its physically motivated anisotropic dependence structure is based on geometric objects moving through space-time according to a velocity vector. We demonstrate that inference based on weighted pairwise likelihood is fast and accurate. The usefulness of our model is illustrated by an application to hourly precipitation data from a study region in Southern France, where it clearly improves on an alternative censored Gaussian space-time random field model. While classical limit models based on threshold-stability fail to appropriately capture relatively fast joint tail decay rates between asymptotic dependence and classical independence, strong empirical evidence from our application and other recent case studies motivates the use of more realistic asymptotic independence models such as ours.

stat.ME

Extremal dependence of random scale constructions

A bivariate random vector can exhibit either asymptotic independence or dependence between the largest values of its components. When used as a statistical model for risk assessment in fields such as finance, insurance or meteorology, it is crucial to understand which of the two asymptotic regimes occurs. Motivated by their ubiquity and flexibility, we consider the extremal dependence properties of vectors with a random scale construction $(X_1,X_2)=R(W_1,W_2)$, with non-degenerate $R>0$ independent of $(W_1,W_2)$. Focusing on the presence and strength of asymptotic tail dependence, as expressed through commonly-used summary parameters, broad factors that affect the results are: the heaviness of the tails of $R$ and $(W_1,W_2)$, the shape of the support of $(W_1,W_2)$, and dependence between $(W_1,W_2)$. When $R$ is distinctly lighter tailed than $(W_1,W_2)$, the extremal dependence of $(X_1,X_2)$ is typically the same as that of $(W_1,W_2)$, whereas similar or heavier tails for $R$ compared to $(W_1,W_2)$ typically result in increased extremal dependence. Similar tail heavinesses represent the most interesting and technical cases, and we find both asymptotic independence and dependence of $(X_1,X_2)$ possible in such cases when $(W_1,W_2)$ exhibit asymptotic independence. The bivariate case often directly extends to higher-dimensional vectors and spatial processes, where the dependence is mainly analyzed in terms of summaries of bivariate sub-vectors. The results unify and extend many existing examples, and we use them to propose new models that encompass both dependence classes.

math.PR

Exceedance-based nonlinear regression of tail dependence

The probability and structure of co-occurrences of extreme values in multivariate data may critically depend on auxiliary information provided by covariates. In this contribution, we develop a flexible generalized additive modeling framework based on high threshold exceedances for estimating covariate-dependent joint tail characteristics for regimes of asymptotic dependence and asymptotic independence. The framework is based on suitably defined marginal pretransformations and projections of the random vector along the directions of the unit simplex, which lead to convenient univariate representations of multivariate exceedances based on the exponential distribution. Good performance of our estimators of a nonparametrically designed influence of covariates on extremal coefficients and tail dependence coefficients are shown through a simulation study. We illustrate the usefulness of our modeling framework on a large dataset of nitrogen dioxide measurements recorded in France between 1999 and 2012, where we use the generalized additive framework for modeling marginal distributions and tail dependence in monthly maxima. Our results imply asymptotic independence of data observed at different stations, and we find that the estimated coefficients of tail dependence decrease as a function of spatial distance and show distinct patterns for different years and for different types of stations (traffic vs. background).

stat.ME

INLA goes extreme: Bayesian tail regression for the estimation of high spatio-temporal quantiles

This work has been motivated by the challenge of the 2017 conference on Extreme-Value Analysis (EVA2017), with the goal of predicting daily precipitation quantiles at the $99.8\%$ level for each month at observed and unobserved locations. We here develop a Bayesian generalized additive modeling framework tailored to estimate complex trends in marginal extremes observed over space and time. Our approach is based on a set of regression equations linked to the exceedance probability above a high threshold and to the size of the excess, the latter being modeled using the generalized Pareto (GP) distribution suggested by Extreme-Value Theory. Latent random effects are modeled additively and semi-parametrically using Gaussian process priors, which provides high flexibility and interpretability. Fast and accurate estimation of posterior distributions may be performed thanks to the Integrated Nested Laplace approximation (INLA), efficiently implemented in the R-INLA software, which we also use for determining a nonstationary threshold based on a model for the body of the distribution. We show that the GP distribution meets the theoretical requirements of INLA, and we then develop a penalized complexity prior specification for the tail index, which is a crucial parameter for extrapolating tail event probabilities. This prior concentrates mass close to a light exponential tail while allowing heavier tails by penalizing the distance to the exponential distribution. We illustrate this methodology through the modeling of spatial and seasonal trends in daily precipitation data provided by the EVA2017 challenge. Capitalizing on R-INLA's fast computation capacities and large distributed computing resources, we conduct an extensive cross-validation study to select model parameters governing the smoothness of trends. Our results outperform simple benchmarks and are comparable to the best-scoring approach.

stat.ME

Spatial random field models based on Lévy indicator convolutions

Process convolutions yield random fields with flexible marginal distributions and dependence beyond Gaussianity, but statistical inference is often hampered by a lack of closed-form marginal distributions, and simulation-based inference may be prohibitively computer-intensive. We here remedy such issues through a class of process convolutions based on smoothing a (d+1)-dimensional Lévy basis with an indicator function kernel to construct a d-dimensional convolution process. Indicator kernels ensure univariate distributions in the Lévy basis family, which provides a sound basis for interpretation, parametric modeling and statistical estimation. We propose a class of isotropic stationary convolution processes constructed through hypograph indicator sets defined as the space between the curve (s,H(s)) of a spherical probability density function H and the plane (s,0). If H is radially decreasing, the covariance is expressed through the univariate distribution function of H. The bivariate joint tail behavior in such convolution processes is studied in detail. Simulation and modeling extensions beyond isotropic stationary spatial models are discussed, including latent process models. For statistical inference of parametric models, we develop pairwise likelihood techniques and illustrate these on spatially indexed weed counts in the Bjertop data set, and on daily wind speed maxima observed over 30 stations in the Netherlands.

stat.ME

Point process-based modeling of multiple debris flow landslides using INLA: an application to the 2009 Messina disaster

We develop a stochastic modeling approach based on spatial point processes of log-Gaussian Cox type for a collection of around 5000 landslide events provoked by a precipitation trigger in Sicily, Italy. Through the embedding into a hierarchical Bayesian estimation framework, we can use the Integrated Nested Laplace Approximation methodology to make inference and obtain the posterior estimates. Several mapping units are useful to partition a given study area in landslide prediction studies. These units hierarchically subdivide the geographic space from the highest grid-based resolution to the stronger morphodynamic-oriented slope units. Here we integrate both mapping units into a single hierarchical model, by treating the landslide triggering locations as a random point pattern. This approach diverges fundamentally from the unanimously used presence-absence structure for areal units since we focus on modeling the expected landslide count jointly within the two mapping units. Predicting this landslide intensity provides more detailed and complete information as compared to the classically used susceptibility mapping approach based on relative probabilities. To illustrate the model's versatility, we compute absolute probability maps of landslide occurrences and check its predictive power over space. While the landslide community typically produces spatial predictive models for landslides only in the sense that covariates are spatially distributed, no actual spatial dependence has been explicitly integrated so far for landslide susceptibility. Our novel approach features a spatial latent effect defined at the slope unit level, allowing us to assess the spatial influence that remains unexplained by the covariates in the model.

stat.AP

Latent Gaussian modeling and INLA: A review with focus on space-time applications

Bayesian hierarchical models with latent Gaussian layers have proven very flexible in capturing complex stochastic behavior and hierarchical structures in high-dimensional spatial and spatio-temporal data. Whereas simulation-based Bayesian inference through Markov Chain Monte Carlo may be hampered by slow convergence and numerical instabilities, the inferential framework of Integrated Nested Laplace Approximation (INLA) is capable to provide accurate and relatively fast analytical approximations to posterior quantities of interest. It heavily relies on the use of Gauss-Markov dependence structures to avoid the numerical bottleneck of high-dimensional nonsparse matrix computations. With a view towards space-time applications, we here review the principal theoretical concepts, model classes and inference tools within the INLA framework. Important elements to construct space-time models are certain spatial Matérn-like Gauss-Markov random fields, obtained as approximate solutions to a stochastic partial differential equation. Efficient implementation of statistical inference tools for a large variety of models is available through the INLA package of the R software. To showcase the practical use of R-INLA and to illustrate its principal commands and syntax, a comprehensive simulation experiment is presented using simulated non Gaussian space-time count data with a first-order autoregressive dependence structure in time.

stat.ME

Bridging Asymptotic Independence and Dependence in Spatial Extremes Using Gaussian Scale Mixtures

Gaussian scale mixtures are constructed as Gaussian processes with a random variance. They have non-Gaussian marginals and can exhibit asymptotic dependence unlike Gaussian processes, which are asymptotically independent except in the case of perfect dependence. In this paper, we study in detail the extremal dependence properties of Gaussian scale mixtures and we unify and extend general results on their joint tail decay rates in both asymptotic dependence and independence cases. Motivated by the analysis of spatial extremes, we propose several flexible yet parsimonious parametric copula models that smoothly interpolate from asymptotic dependence to independence and include the Gaussian dependence as a special case. We show how these new models can be fitted to high threshold exceedances using a censored likelihood approach, and we demonstrate that they provide valuable information about tail characteristics. Our parametric approach outperforms the widely used nonparametric $χ$ and $\barχ$ statistics often used to guide model choice at an exploratory stage by borrowing strength across locations for better estimation of the asymptotic dependence class. We demonstrate the capacity of our methodology by adequately capturing the extremal properties of wind speed data collected in the Pacific Northwest, US.

stat.ME