SearcharxivSearch

arXiv subjects

Emma S. Simpson

Publications and source records attributed to Emma S. Simpson.

9 recordsLinked to original sources

Spatial Extremes at Scale: A Case Study of Surface Skin Temperature and Heat Risk in the United States

Understanding and mapping extreme heat is critical for risk management and public health planning, particularly in regions with complex terrain and heterogeneous climate. We present a case study of extreme heat in the Four Corners region of the United States, using high-resolution surface skin temperature data from the North American Land Data Assimilation System to characterize spatially heterogeneous and seasonally varying extremes across complex terrain, and to assess their implications for heat-related public health risks. Spatial extremes exhibit complex dependencies across geographic regions, which require sophisticated statistical models to capture. While recent advances in spatial extreme value modeling provide flexible representations of joint tail dependencies, statistical inference remains computationally demanding, especially for datasets with a large number of locations. To address this, we propose a random scale mixture process that facilitates Bayesian inference of spatial extremes, and develop scalable inference strategies that leverage advances in spatial modeling and amortized learning. We evaluate the proposed inference methods through large-scale simulation studies, representing the first such extensive study in spatial extremes, and a high-resolution surface skin temperature application in the Four Corners region. Surface skin temperature is particularly useful as a predictor for air temperature, for studying heatwaves and related environmental phenomena, and to calculate heat indices reflecting downstream health risks at any location. Our findings provide insights into efficient, data-driven approaches for modeling spatial extremes, and serve as guidelines for practitioners in the fields of climate science, environmental risk assessment, and beyond.

stat.AP

Accounting for missing data when modelling block maxima

Modelling block maxima using the generalised extreme value (GEV) distribution is a classical and widely used method for studying univariate extremes. It allows for theoretically motivated estimation of return levels, including extrapolation beyond the range of observed data. A frequently overlooked challenge in applying this methodology comes from handling datasets containing missing values. In this case, one cannot be sure whether the true maximum has been recorded in each block, and simply ignoring the issue can lead to biased parameter estimators and, crucially, underestimated return levels. We propose an extension of the standard block maxima approach to overcome such missing data issues. This is achieved by explicitly accounting for the proportion of missing values in each block within the GEV model. Inference is carried out using likelihood-based techniques, and we propose an update to commonly used diagnostic plots to assess model fit. We assess the performance of our method via a simulation study, with results that are competitive with the "ideal" case of having no missing values. The practical use of our methodology is demonstrated on sea surge data from Brest, France, and air pollution data from Plymouth, U.K.

stat.ME

Generative Machine Learning for Multivariate Angular Simulation

With the recent development of new geometric and angular-radial frameworks for multivariate extremes, reliably simulating from angular variables in moderate-to-high dimensions is of increasing importance. Empirical approaches have the benefit of simplicity, and work reasonably well in low dimensions, but as the number of variables increases, they can lack the required flexibility and scalability. Classical parametric models for angular variables, such as the von Mises--Fisher distribution (vMF), provide an alternative. Exploiting finite mixtures of vMF distributions increases their flexibility, but there are cases where, without letting the number of mixture components grow considerably, a mixture model with a fixed number of components is not sufficient to capture the intricate features that can arise in data. Owing to their flexibility, generative deep learning methods are able to capture complex data structures; they therefore have the potential to be useful in the simulation of multivariate angular variables. In this paper, we introduce a range of deep learning approaches for this task, including generative adversarial networks, normalizing flows and flow matching. We assess their performance via a range of metrics, and make comparisons to the more classical approach of using a finite mixture of vMF distributions. The methods are also applied to a metocean data set, with diagnostics indicating strong performance, demonstrating the applicability of such techniques to real-world, complex data structures.

stat.ML

Spatial extremal modelling: A case study on the interplay between margins and dependence

It is no secret that statistical modelling often involves making simplifying assumptions when attempting to study complex stochastic phenomena. Spatial modelling of extreme values is no exception, with one of the most common such assumptions being stationarity in the marginal and/or dependence features. If non-stationarity has been detected in the marginal distributions, it is tempting to try to model this while assuming stationarity in the dependence, without necessarily putting this latter assumption through thorough testing. However, margins and dependence are often intricately connected and the detection of non-stationarity in one feature might affect the detection of non-stationarity in the other. This work is an in-depth case study of this interrelationship, with a particular focus on a spatio-temporal environmental application exhibiting well-documented marginal non-stationarity. Specifically, we compare and contrast four different marginal detrending approaches in terms of our post-detrending ability to detect temporal non-stationarity in the spatial extremal dependence structure of a sea surface temperature dataset from the Red Sea.

stat.AP

Estimating the limiting shape of bivariate scaled sample clouds: with additional benefits of self-consistent inference for existing extremal dependence properties

The key to successful statistical analysis of bivariate extreme events lies in flexible modelling of the tail dependence relationship between the two variables. In the extreme value theory literature, various techniques are available to model separate aspects of tail dependence, based on different asymptotic limits. Results from Balkema and Nolde (2010) and Nolde (2014) highlight the importance of studying the limiting shape of an appropriately-scaled sample cloud when characterising the whole joint tail. We now develop the first statistical inference for this limit set, which has considerable practical importance for a unified inference framework across different aspects of the joint tail. Moreover, Nolde and Wadsworth (2022) link this limit set to various existing extremal dependence frameworks. Hence, a by-product of our new limit set inference is the first set of self-consistent estimators for several extremal dependence measures, avoiding the current possibility of contradictory conclusions. In simulations, our limit set estimator is successful across a range of distributions, and the corresponding extremal dependence estimators provide a major joint improvement and small marginal improvements over existing techniques. We consider an application to sea wave heights, where our estimates successfully capture the expected weakening extremal dependence as the distance between locations increases.

stat.ME

A geometric investigation into the tail dependence of vine copulas

Vine copulas are a type of multivariate dependence model, composed of a collection of bivariate copulas that are combined according to a specific underlying graphical structure. Their flexibility and practicality in moderate and high dimensions have contributed to the popularity of vine copulas, but relatively little attention has been paid to their extremal properties. To address this issue, we present results on the tail dependence properties of some of the most widely studied vine copula classes. We focus our study on the coefficient of tail dependence and the asymptotic shape of the sample cloud, which we calculate using the geometric approach of Nolde (2014). We offer new insights by presenting results for trivariate vine copulas constructed from asymptotically dependent and asymptotically independent bivariate copulas, focusing on bivariate extreme value and inverted extreme value copulas, with additional detail provided for logistic and inverted logistic examples. We also present new theory for a class of higher dimensional vine copulas, constructed from bivariate inverted extreme value copulas.

math.ST

High-dimensional modeling of spatial and spatio-temporal conditional extremes using INLA and Gaussian Markov random fields

The conditional extremes framework allows for event-based stochastic modeling of dependent extremes, and has recently been extended to spatial and spatio-temporal settings. After standardizing the marginal distributions and applying an appropriate linear normalization, certain non-stationary Gaussian processes can be used as asymptotically-motivated models for the process conditioned on threshold exceedances at a fixed reference location and time. In this work, we adapt existing conditional extremes models to allow for the handling of large spatial datasets. This involves specifying the model for spatial observations at $d$ locations in terms of a latent $m\ll d$ dimensional Gaussian model, whose structure is specified by a Gaussian Markov random field. We perform Bayesian inference for such models for datasets containing thousands of observation locations using the integrated nested Laplace approximation, or INLA. We explain how constraints on the spatial and spatio-temporal Gaussian processes, arising from the conditioning mechanism, can be implemented through the latent variable approach without losing the computationally convenient Markov property. We discuss tools for the comparison of models via their posterior distributions, and illustrate the flexibility of the approach with gridded Red Sea surface temperature data at over $6,000$ observed locations. Posterior sampling is exploited to study the probability distribution of cluster functionals of spatial and spatio-temporal extreme episodes.

stat.ME

Conditional Modelling of Spatio-Temporal Extremes for Red Sea Surface Temperatures

Recent extreme value theory literature has seen significant emphasis on the modelling of spatial extremes, with comparatively little consideration of spatio-temporal extensions. This neglects an important feature of extreme events: their evolution over time. Many existing models for the spatial case are limited by the number of locations they can handle; this impedes extension to space-time settings, where models for higher dimensions are required. Moreover, the spatio-temporal models that do exist are restrictive in terms of the range of extremal dependence types they can capture. Recently, conditional approaches for studying multivariate and spatial extremes have been proposed, which enjoy benefits in terms of computational efficiency and an ability to capture both asymptotic dependence and asymptotic independence. We extend this class of models to a spatio-temporal setting, conditioning on the occurrence of an extreme value at a single space-time location. We adopt a composite likelihood approach for inference, which combines information from full likelihoods across multiple space-time conditioning locations. We apply our model to Red Sea surface temperatures, show that it fits well using a range of diagnostic plots, and demonstrate how it can be used to assess the risk of coral bleaching attributed to high water temperatures over consecutive days.

stat.ME

Determining the Dependence Structure of Multivariate Extremes

In multivariate extreme value analysis, the nature of the extremal dependence between variables should be considered when selecting appropriate statistical models. Interest often lies with determining which subsets of variables can take their largest values simultaneously, while the others are of smaller order. Our approach to this problem exploits hidden regular variation properties on a collection of non-standard cones and provides a new set of indices that reveal aspects of the extremal dependence structure not available through existing measures of dependence. We derive theoretical properties of these indices, demonstrate their value through a series of examples, and develop methods of inference that also estimate the proportion of extremal mass associated with each cone. We apply the methods to UK river flows, estimating the probabilities of different subsets of sites being large simultaneously.

stat.ME