SearcharxivSearch

arXiv subjects

Jonathan A. Tawn

Publications and source records attributed to Jonathan A. Tawn.

At least 19 recordsLinked to original sources

Spatio-temporal analysis of extreme winter temperatures in Ireland

We analyse extreme daily minimum temperatures in winter months over the island of Ireland from 1950-2022. We model the marginal distributions of extreme winter minima using a generalised Pareto distribution (GPD), capturing temporal and spatial non-stationarities in the parameters of the GPD. We investigate two independent temporal non-stationarities in extreme winter minima. We model the long-term trend in magnitude of extreme winter minima as well as short-term, large fluctuations in magnitude caused by anomalous behaviour of the jet stream. We measure magnitudes of spatial events with a carefully chosen risk function and fit an r-Pareto process to extreme events exceeding a high-risk threshold. Our analysis is based on synoptic data observations courtesy of Met Éireann and the Met Office. We show that the frequency of extreme cold winter events is decreasing over the study period. The magnitude of extreme winter events is also decreasing, indicating that winters are warming, and apparently warming at a faster rate than extreme summer temperatures. We also show that extremely cold winter temperatures are warming at a faster rate than non-extreme winter temperatures. We find that a climate model output previously shown to be informative as a covariate for modelling extremely warm summer temperatures is less effective as a covariate for extremely cold winter temperatures. However, we show that the climate model is useful for informing a non-extreme temperature model.

physics.ao-ph

Gaussian mixture copulas for flexible dependence modelling in the body and tails of joint distributions

Fully describing the entire data set is essential in multivariate risk assessment, since moderate levels of one variable can influence another, potentially leading it to be extreme. Additionally, modelling both non-extreme and extreme events within a single framework avoids the need to select a threshold vector used to determine an extremal region, or the requirement to add flexibility to bridge between separate models for the body and tail regions. We propose a copula model, based on a mixture of Gaussian distributions, as this model avoids the need to define an extremal region, it is scalable to dimensions beyond the bivariate case, and it can handle both asymptotic dependent and asymptotic independent extremal dependence structures. We apply the proposed model through simulations and to a 5-dimensional seasonal air pollution data set, previously analysed in the multivariate extremes literature. Through pairwise, trivariate and 5-dimensional analyses, we show the flexibility of the Gaussian mixture copula in capturing different joint distributional behaviours and its ability to identify potential graphical structure features, both of which can vary across the body and tail regions.

stat.ME

Automated threshold selection and associated inference uncertainty for univariate extremes

Threshold selection is a fundamental problem in any threshold-based extreme value analysis. While models are asymptotically motivated, selecting an appropriate threshold for finite samples is difficult and highly subjective through standard methods. Inference for high quantiles can also be highly sensitive to the choice of threshold. Too low a threshold choice leads to bias in the fit of the extreme value model, while too high a choice leads to unnecessary additional uncertainty in the estimation of model parameters. We develop a novel methodology for automated threshold selection that directly tackles this bias-variance trade-off. We also develop a method to account for the uncertainty in the threshold estimation and propagate this uncertainty through to high quantile inference. Through a simulation study, we demonstrate the effectiveness of our method for threshold selection and subsequent extreme quantile estimation, relative to the leading existing methods, and show how the method's effectiveness is not sensitive to the tuning parameters. We apply our method to the well-known, troublesome example of the River Nidd dataset.

stat.ME

Estimating the limiting shape of bivariate scaled sample clouds: with additional benefits of self-consistent inference for existing extremal dependence properties

The key to successful statistical analysis of bivariate extreme events lies in flexible modelling of the tail dependence relationship between the two variables. In the extreme value theory literature, various techniques are available to model separate aspects of tail dependence, based on different asymptotic limits. Results from Balkema and Nolde (2010) and Nolde (2014) highlight the importance of studying the limiting shape of an appropriately-scaled sample cloud when characterising the whole joint tail. We now develop the first statistical inference for this limit set, which has considerable practical importance for a unified inference framework across different aspects of the joint tail. Moreover, Nolde and Wadsworth (2022) link this limit set to various existing extremal dependence frameworks. Hence, a by-product of our new limit set inference is the first set of self-consistent estimators for several extremal dependence measures, avoiding the current possibility of contradictory conclusions. In simulations, our limit set estimator is successful across a range of distributions, and the corresponding extremal dependence estimators provide a major joint improvement and small marginal improvements over existing techniques. We consider an application to sea wave heights, where our estimates successfully capture the expected weakening extremal dependence as the distance between locations increases.

stat.ME

Hidden tail chains and recurrence equations for dependence parameters associated with extremes of higher-order Markov chains

We derive some key extremal features for $k$th order Markov chains that can be used to understand how the process moves between an extreme state and the body of the process. The chains are studied given that there is an exceedance of a threshold, as the threshold tends to the upper endpoint of the distribution. Unlike previous studies with $k>1$, we consider processes where standard limit theory describes each extreme event as a single observation without any information about the transition to and from the body of the distribution. Our work uses different asymptotic theory which results in non-degenerate limit laws for such processes. We study the extremal properties of the initial distribution and the transition probability kernel of the Markov chain under weak assumptions for broad classes of extremal dependence structures that cover both asymptotically dependent and asymptotically independent Markov chains. For chains with $k>1$, the transition of the chain away from the exceedance involves novel functions of the $k$ previous states, in comparison to just the single value, when $k=1$. This leads to an increase in the complexity of determining the form of this class of functions, their properties and the method of their derivation in applications. We find that it is possible to derive an affine normalization, dependent on the threshold excess, such that non-degenerate limiting behaviour of the process is assured for all lags. These normalization functions have an attractive structure that has parallels to the Yule-Walker equations. Furthermore, the limiting process is always linear in the innovations. We illustrate the results with the study of $k$th order stationary Markov chains with exponential margins based on widely studied families of copula dependence structures.

math.ST

Joint Estimation of Extreme Spatially Aggregated Precipitation at Different Scales through Mixture Modelling

Although most models for rainfall extremes focus on point-wise values, it is aggregated precipitation over areas up to river catchment scale that is of the most interest. To capture the joint behaviour of precipitation aggregates evaluated at different spatial scales, parsimonious and effective models must be built with knowledge of the underlying spatial process. Precipitation is driven by a mixture of processes acting at different scales and intensities, e.g., convective and frontal, with extremes of aggregates for typical catchment sizes arising from extremes of only one of these processes, rather than a combination of them. High-intensity convective events cause extreme spatial aggregates at small scales but the contribution of lower-intensity large-scale fronts is likely to increase as the area aggregated increases. Thus, to capture small to large scale spatial aggregates within a single approach requires a model that can accurately capture the extremal properties of both convective and frontal events. Previous extreme value methods have ignored this mixture structure; we propose a spatial extreme value model which is a mixture of two components with different marginal and dependence models that are able to capture the extremal behaviour of convective and frontal rainfall and more faithfully reproduces spatial aggregates for a wide range of scales. Modelling extremes of the frontal component raises new challenges due to it exhibiting strong long-range extremal spatial dependence. Our modelling approach is applied to fine-scale, high-dimensional, gridded precipitation data. We show that accounting for the mixture structure improves the joint inference on extremes of spatial aggregates over regions of different sizes.

stat.AP

Accounting for Climate Change in Extreme Sea Level Estimation

Extreme sea level estimates are fundamental for mitigating against coastal flooding as they provide insight for defence engineering. As the global climate changes, rising sea levels combined with increases in storm intensity and frequency pose an increasing risk to coastline communities. We present a new method for estimating extreme sea levels that accounts for the effects of climate change on extreme events that are not accounted for by mean sea level trends. We follow a joint probabilities methodology, considering skew surge and peak tides as the only components of sea levels. We model extreme skew surges using a non-stationary generalised Pareto distribution (GPD) with covariates accounting for climate change, seasonality and skew surge-peak tide interaction. We develop methods to efficiently test for extreme skew surge trends across different coastlines and seasons. We illustrate our methods using data from four UK tide gauges.

stat.ME

Accounting for Seasonality in Extreme Sea Level Estimation

Reliable estimates of sea level return levels are crucial for coastal flooding risk assessments and for coastal flood defence design. We describe a novel method for estimating extreme sea levels that is the first to capture seasonality, interannual variations and longer term changes. We use a joint probabilities method, with skew surge and peak tide as two sea level components. The tidal regime is predictable but skew surges are stochastic. We present a statistical model for skew surges, where the main body of the distribution is modelled empirically whilst a non-stationary generalised Pareto distribution (GPD) is used for the upper tail. We capture within-year seasonality by introducing a daily covariate to the GPD model and allowing the distribution of peak tides to change over months and years. Skew surge-peak tide dependence is accounted for via a tidal covariate in the GPD model and we adjust for skew surge temporal dependence through the subasymptotic extremal index. We incorporate spatial prior information in our GPD model to reduce the uncertainty associated with the highest return level estimates. Our results are an improvement on current return level estimates, with previous methods typically underestimating. We illustrate our method at four UK tide gauges.

stat.AP

Modelling Extremes of Spatial Aggregates of Precipitation using Conditional Methods

Inference on the extremal behaviour of spatial aggregates of precipitation is important for quantifying river flood risk. There are two classes of previous approach, with one failing to ensure self-consistency in inference across different regions of aggregation and the other imposing highly restrictive assumptions. To overcome these issues, we propose a model for high-resolution precipitation data, from which we can simulate realistic fields and explore the behaviour of spatial aggregates. Recent developments have seen spatial extensions of the Heffernan and Tawn (2004) model for conditional multivariate extremes, which can handle a wide range of dependence structures. Our contribution is twofold: extensions and improvements of this approach and its model inference for high-dimensional data; and a novel framework for deriving aggregates addressing edge effects and sub-regions without rain. We apply our modelling approach to gridded East-Anglia, UK precipitation data. Return-level curves for spatial aggregates over different regions of various sizes are estimated and shown to fit very well to the data.

stat.ME

On the Tail Behaviour of Aggregated Random Variables

In many areas of interest, modern risk assessment requires estimation of the extremal behaviour of sums of random variables. We derive the first order upper-tail behaviour of the weighted sum of bivariate random variables under weak assumptions on their marginal distributions and their copula. The extremal behaviour of the marginal variables is characterised by the generalised Pareto distribution and their extremal dependence through subclasses of the limiting representations of Ledford and Tawn (1997) and Heffernan and Tawn (2004). We find that the upper tail behaviour of the aggregate is driven by different factors dependent on the signs of the marginal shape parameters; if they are both negative, the extremal behaviour of the aggregate is determined by both marginal shape parameters and the coefficient of asymptotic independence (Ledford and Tawn, 1996); if they are both positive or have different signs, the upper-tail behaviour of the aggregate is given solely by the largest marginal shape. We also derive the aggregate upper-tail behaviour for some well known copulae which reveals further insight into the tail structure when the copula falls outside the conditions for the subclasses of the limiting dependence representations.

math.ST

Inference for extreme earthquake magnitudes accounting for a time-varying measurement process

Investment in measuring a process more completely or accurately is only useful if these improvements can be utilised during modelling and inference. We consider how improvements to data quality over time can be incorporated when selecting a modelling threshold and in the subsequent inference of an extreme value analysis. Motivated by earthquake catalogues, we consider variable data quality in the form of rounded and incompletely observed data. We develop an approach to select a time-varying modelling threshold that makes best use of the available data, accounting for uncertainty in the magnitude model and for the rounding of observations. We show the benefits of the proposed approach on simulated data and apply the method to a catalogue of earthquakes induced by gas extraction in the Netherlands. This more than doubles the usable catalogue size and greatly increases the precision of high magnitude quantile estimates. This has important consequences for the design and cost of earthquake defences. For the first time, we find compelling data-driven evidence against the applicability of the Gutenberg-Richer law to these earthquakes. Furthermore, our approach to automated threshold selection appears to have much potential for generic applications of extreme value methods.

stat.ME

A geometric investigation into the tail dependence of vine copulas

Vine copulas are a type of multivariate dependence model, composed of a collection of bivariate copulas that are combined according to a specific underlying graphical structure. Their flexibility and practicality in moderate and high dimensions have contributed to the popularity of vine copulas, but relatively little attention has been paid to their extremal properties. To address this issue, we present results on the tail dependence properties of some of the most widely studied vine copula classes. We focus our study on the coefficient of tail dependence and the asymptotic shape of the sample cloud, which we calculate using the geometric approach of Nolde (2014). We offer new insights by presenting results for trivariate vine copulas constructed from asymptotically dependent and asymptotically independent bivariate copulas, focusing on bivariate extreme value and inverted extreme value copulas, with additional detail provided for logistic and inverted logistic examples. We also present new theory for a class of higher dimensional vine copulas, constructed from bivariate inverted extreme value copulas.

math.ST

Determining the Dependence Structure of Multivariate Extremes

In multivariate extreme value analysis, the nature of the extremal dependence between variables should be considered when selecting appropriate statistical models. Interest often lies with determining which subsets of variables can take their largest values simultaneously, while the others are of smaller order. Our approach to this problem exploits hidden regular variation properties on a collection of non-standard cones and provides a new set of indices that reveal aspects of the extremal dependence structure not available through existing measures of dependence. We derive theoretical properties of these indices, demonstrate their value through a series of examples, and develop methods of inference that also estimate the proportion of extremal mass associated with each cone. We apply the methods to UK river flows, estimating the probabilities of different subsets of sites being large simultaneously.

stat.ME

Evaluation of extremal properties of GARCH(p,q) processes

Generalized autoregressive conditionally heteroskedastic (GARCH) processes are widely used for modelling features commonly found in observed financial returns. The extremal properties of these processes are of considerable interest for market risk management. For the simplest GARCH(p,q) process, with max(p,q) = 1, all extremal features have been fully characterised. Although the marginal features of extreme values of the process have been theoretically characterised when max(p, q) >= 2, much remains to be found about both marginal and dependence structure during extreme excursions. Specifically, a reliable method is required for evaluating the tail index, which regulates the marginal tail behaviour and there is a need for methods and algorithms for determining clustering. In particular, for the latter, the mean number of extreme values in a short-term cluster, i.e., the reciprocal of the extremal index, has only been characterised in special cases which exclude all GARCH(p,q) processes that are used in practice. Although recent research has identified the multivariate regular variation property of stationary GARCH(p,q) processes, currently there are no reliable methods for numerically evaluating key components of these characterisations. We overcome these issues and are able to generate the forward tail chain of the process to derive the extremal index and a range of other cluster functionals for all GARCH(p, q) processes including integrated GARCH processes and processes with unbounded and asymmetric innovations. The new theory and methods we present extend to assessing the strict stationarity and extremal properties for a much broader class of stochastic recurrence equations.

stat.CO

Modelling the spatial extent and severity of extreme European windstorms

Windstorms are a primary natural hazard affecting Europe that are commonly linked to substantial property and infrastructural damage and are responsible for the largest spatially aggregated financial losses. Such extreme winds are typically generated by extratropical cyclone systems originating in the North Atlantic and passing over Europe. Previous statistical studies tend to model extreme winds at a given set of sites, corresponding to inference in a Eulerian framework. Such inference cannot incorporate knowledge of the life cycle and progression of extratropical cyclones across the region and is forced to make restrictive assumptions about the extremal dependence structure. We take an entirely different approach which overcomes these limitations by working in a Lagrangian framework. Specifically, we model the development of windstorms over time, preserving the physical characteristics linking the windstorm and the cyclone track, the path of local vorticity maxima, and make a key finding that the spatial extent of extratropical windstorms becomes more localised as its magnitude increases irrespective of the location of the storm track. Our model allows simulation of synthetic windstorm events to derive the joint distributional features over any set of sites giving physically consistent extrapolations to rarer events. From such simulations improved estimates of this hazard can be achieved both in terms of intensity and area affected.

stat.AP

A stochastic model for the lifecycle and track of extreme extratropical cyclones in the North Atlantic

Extratropical cyclones are large-scale weather systems which are often the source of extreme weather events in Northern Europe, often leading to mass infrastructural damage and casualties. Such systems create a local vorticity maxima which tracks across the Atlantic Ocean and from which can be determined a climatology for the region. While there have been considerable advances in developing algorithms for extracting the track and evolution of cyclones from reanalysis datasets, the data record is relatively short. This justifies the need for a statistical model to represent the more extreme characteristics of these weather systems, specifically their intensity and the spatial variability in their tracks. This paper presents a novel simulation-based approach to modelling the lifecycle of extratropical cyclones in terms of both their tracks and vorticity, incorporating various aspects of cyclone evolution and movement. By drawing on methods from extreme value analysis, we can simulate more extreme storms than those observed, representing a useful tool for practitioners concerned with risk assessment with regard to these weather systems.

stat.AP

Penultimate Analysis of the Conditional Multivariate Extremes Tail Model

Models for extreme values are generally derived from limit results, which are meant to be good enough approximations when applied to finite samples. Depending on the speed of convergence of the process underlying the data, these approximations may fail to represent subasymptotic features present in the data, and thus may introduce bias. The case of univariate maxima has been widely explored in the literature, a prominent example being the slow convergence to their Gumbel limit of Gaussian maxima, which are better approximated by a negative Weibull distribution at finite levels. In the context of subasymptotic multivariate extremes, research has only dealt with specific cases related to componentwise maxima and multivariate regular variation. This paper explores the conditional extremes model (Heffernan and Tawn, 2004) in order to shed light on its finite-sample behaviour and to reduce the bias of extrapolations beyond the range of the available data. We identify second-order features for different types of conditional copulas, and obtain results that echo those from the univariate context. These results suggest possible extensions of the conditional tail model, which will enable it to be fitted at less extreme thresholds.

math.ST

A Poisson process reparameterisation for Bayesian inference for extremes

A common approach to modelling extreme values is to consider the excesses above a high threshold as realisations of a non-homogeneous Poisson process. While this method offers the advantage of modelling using threshold-invariant extreme value parameters, the dependence between these parameters makes estimation more difficult. We present a novel approach for Bayesian estimation of the Poisson process model parameters by reparameterising in terms of a tuning parameter $m$. This paper presents a method for choosing the optimal value of m that near-orthogonalises the parameters, which is achieved by minimising the correlation between the asymptotic posterior distribution of the parameters. This choice of m ensures more rapid convergence and efficient sampling from the joint posterior distribution using Markov Chain Monte Carlo methods. Samples from the parameterisation of interest are then obtained by a simple transform. Results are presented in the cases of identically and non-identically distributed models for extreme rainfall in Cumbria, UK.

stat.AP