SearcharxivSearch

arXiv subjects

Jonathan Tawn

Publications and source records attributed to Jonathan Tawn.

12 recordsLinked to original sources

A framework for statistical modelling of the extremes of longitudinal data, applied to elite swimming

We develop methods, based on extreme value theory, for analysing observations in the tails of longitudinal data, i.e., a data set consisting of a large number of short time series, which are typically irregularly and non-simultaneously sampled, yet have some commonality in the structure of each series and exhibit independence between time series. Extreme value theory has not been considered previously for the unique features of longitudinal data. Across time series the data are assumed to follow a common generalised Pareto distribution, above a high threshold. To account for temporal dependence of such data we require a model to describe (i) the variation between the different time series properties, (ii) the changes in distribution over time, and (iii) the temporal dependence within each series. Our methodology has the flexibility to capture both asymptotic dependence and asymptotic independence, with this characteristic determined by the data. Bayesian inference is used given the need for inference of parameters that are unique to each time series. Our novel methodology is illustrated through the analysis of data from elite swimmers in the men's 100m breaststroke. Unlike previous analyses of personal-best data in this event, we are able to make inference about the careers of individual swimmers - such as the probability an individual will break the world record or swim the fastest time next year.

stat.ME

Temporal evolution of the extreme excursions of multivariate $k$th order Markov processes with application to oceanographic data

We develop two models for the temporal evolution of extreme events of multivariate $k$th order Markov processes. The foundation of our methodology lies in the conditional extremes model of Heffernan & Tawn (2004), and it naturally extends the work of Winter & Tawn (2016,2017) and Tendijck et al. (2019) to include multivariate random variables. We use cross-validation-type techniques to develop a model order selection procedure, and we test our models on two-dimensional meteorological-oceanographic data with directional covariates for a location in the northern North Sea. We conclude that the newly-developed models perform better than the widely used historical matching methodology for these data.

stat.ME

Extremal Characteristics of Conditional Models

Conditionally specified models are often used to describe complex multivariate data. Such models assume implicit structures on the extremes. So far, no methodology exists for calculating extremal characteristics of conditional models since the copula and marginals are not expressed in closed forms. We consider bivariate conditional models that specify the distribution of $X$ and the distribution of $Y$ conditional on $X$. We provide tools to quantify implicit assumptions on the extremes of this class of models. In particular, these tools allow us to approximate the distribution of the tail of $Y$ and the coefficient of asymptotic independence $\eta$ in closed forms. We apply these methods to a widely used conditional model for wave height and wave period. Moreover, we introduce a new condition on the parameter space for the conditional extremes model of Heffernan and Tawn (2004), and prove that the conditional extremes model does not capture $\eta$, when $\eta<1$.

math.ST

Inference for extreme spatial temperature events in a changing climate with application to Ireland

We investigate the changing nature of the frequency, magnitude and spatial extent of extreme temperatures in Ireland from 1931 to 2022. We develop an extreme value model that captures spatial and temporal non-stationarity in extreme daily maximum temperature data. We model the tails of the marginal variables using the generalised Pareto distribution and the spatial dependence of extreme events by a semi-parametric Brown-Resnick r-generalised Pareto process, with parameters of each model allowed to change over time. We use weather station observations for modelling extreme events since data from climate models (not conditioned on observational data) can over-smooth these events and have trends determined by the specific climate model configuration. However, climate models do provide valuable information about the detailed physiography over Ireland and the associated climate response. We propose novel methods which exploit the climate model data to overcome issues linked to the sparse and biased sampling of the observations. Our analysis identifies a temporal change in the marginal behaviour of extreme temperature events over the study domain, which is much larger than the change in mean temperature levels over this time window. We illustrate how these characteristics result in increased spatial coverage of the events that exceed critical temperatures.

stat.ME

Modelling intransitivity in pairwise comparisons with application to baseball data

The seminal Bradley-Terry model exhibits transitivity, i.e., the property that the probabilities of player A beating B and B beating C give the probability of A beating C, with these probabilities determined by a skill parameter for each player. Such transitive models do not account for different strategies of play between each pair of players, which gives rise to {\it intransitivity}. Various intransitive parametric models have been proposed but they lack the flexibility to cover the different strategies across $n$ players, with the $O(n^2)$ values of intransitivity modelled using O(n) parameters, whilst they are not parsimonious when the intransitivity is simple. We overcome their lack of adaptability by allocating each pair of players to one of a random number of $K$ intransitivity levels, each level representing a different strategy. Our novel approach for the skill parameters involves having the $n$ players allocated to a random number of $A<n$ distinct skill levels, to improve efficiency and avoid false rankings. Although we may have to estimate up to $O(n^2)$ unknown parameters for $(A,K)$ we anticipate that in many practical contexts $A+K < n$. Using a Bayesian hierarchical model, $(A,K)$ are treated as unknown, and inference is conducted via a reversible jump Markov chain Monte Carlo (RJMCMC) algorithm. Our semi-parametric model, which gives the Bradley-Terry model when $(A=n-1, K=0)$, is shown to have an improved fit relative to the Bradley-Terry, and the existing intransitivity models, in out-of-sample testing when applied to simulated and American League baseball data. Supplementary materials for the article are available online.

stat.ME

Higher-dimensional spatial extremes via single-site conditioning

Currently available models for spatial extremes suffer either from inflexibility in the dependence structures that they can capture, lack of scalability to high dimensions, or in most cases, both of these. We present an approach to spatial extreme value theory based on the conditional multivariate extreme value model, whereby the limit theory is formed through conditioning upon the value at a particular site being extreme. The ensuing methodology allows for a flexible class of dependence structures, as well as models that can be fitted in high dimensions. To overcome issues of conditioning on a single site, we suggest a joint inference scheme based on all observation locations, and implement an importance sampling algorithm to provide spatial realizations and estimates of quantities conditioning upon the process being extreme at any of one of an arbitrary set of locations. The modelling approach is applied to Australian summer temperature extremes, permitting assessment of the spatial extent of high temperature events over the continent.

stat.ME

Ranking, and other Properties, of Elite Swimmers using Extreme Value Theory

The International Swimming Federation (FINA) uses a very simple points system with the aim to rank swimmers across all Olympic events. The points acquired is a function of the ratio of the recorded time and the current world record for that event. With some world records considered "better" than others however, bias is introduced between events, with some being much harder to attain points where the world record is hard to beat. A model based on extreme value theory will be introduced, where swim-times are modelled through their rate of occurrence, and with the distribution of the best times following a generalised Pareto distribution. Within this framework, the strength of a particular swim is judged based on its position compared to the whole distribution of swim-times, rather than just the world record. This model also accounts for the date of the swim, as training methods improve over the years, as well as changes in technology, such as full body suits. The parameters of the generalised Pareto distribution, for each of the 34 individual Olympic events, will be shown to vary with a covariate, leading to a novel single unified description of swim quality over all events and time. This structure, which allows information to be shared across all strokes, distances, and genders, improves the predictive power as well as the model robustness compared to equivalent independent models. A by-product of the model is that it is possible to estimate other features of interest, such as the ultimate possible time, the distribution of new world records for any event, and to correct swim times for the effect of full body suits. The methods will be illustrated using a dataset of the best 500 swim-times for each event in the period 2001-2018.

stat.AP

Inference for extreme values under threshold-based stopping rules

There is a propensity for an extreme value analyses to be conducted as a consequence of the occurrence of a large flooding event. This timing of the analysis introduces bias and poor coverage probabilities into the associated risk assessments and leads subsequently to inefficient flood protection schemes. We explore these problems through studying stochastic stopping criteria and propose new likelihood-based inferences that mitigate against these difficulties. Our methods are illustrated through the analysis of the river Lune, following it experiencing the UK's largest ever measured flow event in 2015. We show that without accounting for this stopping feature there would be substantial over-design in response to the event.

stat.AP

Modelling the clustering of extreme events for short-term risk assessment

Having reliable estimates of the occurrence rates of extreme events is highly important for insurance companies, government agencies and the general public. The rarity of an extreme event is typically expressed through its return period, i.e., the expected waiting time between events of the observed size if the extreme events of the processes are independent and identically distributed. A major limitation with this measure is when an unexpectedly high number of events occur within the next few months immediately after a \textit{T} year event, with \textit{T} large. Such events undermine the trust in the quality of these risk estimates. The clustering of apparently independent extreme events can occur as a result of local non-stationarity of the process, which can be explained by covariates or random effects. We show how accounting for these covariates and random effects provides more accurate estimates of return levels and aids short-term risk assessment through the use of a new risk measure, which provides evidence of risk which is complementary to the return period.

stat.ME

Model-based inference of conditional extreme value distributions with hydrological applications

Multivariate extreme value models are used to estimate joint risk in a number of applications, with a particular focus on environmental fields ranging from climatology and hydrology to oceanography and seismic hazards. The semi-parametric conditional extreme value model of Heffernan and Tawn (2004) involving a multivariate regression provides the most suitable of current statistical models in terms of its flexibility to handle a range of extremal dependence classes. However, the standard inference for the joint distribution of the residuals of this model suffers from the curse of dimensionality since in a $d$-dimensional application it involves a $d-1$-dimensional non-parametric density estimator, which requires, for accuracy, a number points and commensurate effort that is exponential in $d$. Furthermore, it does not allow for any partially missing observations to be included and a previous proposal to address this is extremely computationally intensive, making its use prohibitive if the proportion of missing data is non-trivial. We propose to replace the $d-1$-dimensional non-parametric density estimator with a model-based copula with univariate marginal densities estimated using kernel methods. This approach provides statistically and computationally efficient estimates whatever the dimension, $d$ or the degree of missing data. Evidence is presented to show that the benefits of this approach substantially outweigh potential mis-specification errors. The methods are illustrated through the analysis of UK river flow data at a network of 46 sites and assessing the rarity of the 2015 floods in north west England.

stat.ME

Modelling across extremal dependence classes

Different dependence scenarios can arise in multivariate extremes, entailing careful selection of an appropriate class of models. In bivariate extremes, the variables are either asymptotically dependent or are asymptotically independent. Most available statistical models suit one or other of these cases, but not both, resulting in a stage in the inference that is unaccounted for, but can substantially impact subsequent extrapolation. Existing modelling solutions to this problem are either applicable only on sub-domains, or appeal to multiple limit theories. We introduce a unified representation for bivariate extremes that encompasses a wide variety of dependence scenarios, and applies when at least one variable is large. Our representation motivates a parametric model that encompasses both dependence classes. We implement a simple version of this model, and show that it performs well in a range of settings.

stat.ME