Searcharxiv⌕ Search

arXiv subjects

Holger Rootzén

Publications and source records attributed to Holger Rootzén.

14 recordsLinked to original sources

Simulation of Multivariate Extremes: a Wasserstein-Aitchison GAN approach

Economically responsible mitigation of multivariate extreme risks-such as extreme rainfall over large areas, large simultaneous variations in many stock prices, or widespread breakdowns in transportation systems-requires assessing the resilience of the systems under plausible stress scenarios. This paper uses Extreme Value Theory (EVT) to develop a new approach to simulating such multivariate extreme events. Specifically, we assume that after transformation to a standard scale the distribution of the random phenomenon of interest is multivariate regular varying and use this to provide a sampling procedure for extremes on the original scale. Our procedure combines a Wasserstein-Aitchison Generative Adversarial Network (WA-GAN) to simulate the tail dependence structure on the standard scale with joint modeling of the univariate marginal tails on the original scale. The WA-GAN procedure relies on the angular measure-encoding the distribution on the unit simplex of the angles of extreme observations-after transformation to Aitchison coordinates, which allows the Wasserstein-GAN algorithm to be run in a linear space. Our method is applied both to simulated data under various tail dependence scenarios and to a financial data set from the Kenneth French Data Library. The proposed algorithm demonstrates strong performance compared to existing alternatives in the literature, both in capturing tail dependence structures and in generating accurate new extreme observations.

stat.ML↗

Fast and robust cross-validation-based scoring rule inference for spatial statistics

Scoring rules are aimed at evaluation of the quality of predictions, but can also be used for estimation of parameters in statistical models. We propose estimating parameters of multivariate spatial models by maximising the average leave-one-out cross-validation score. This method, LOOS, thus optimises predictions instead of maximising the likelihood. The method allows for fast computations for Gaussian models with sparse precision matrices, such as spatial Markov models. It also makes it possible to tailor the estimator's robustness to outliers and their sensitivity to spatial variations of uncertainty through the choice of the scoring rule which is used in the maximisation. The effects of the choice of scoring rule which is used in LOOS are studied by simulation in terms of computation time, statistical efficiency, and robustness. Various popular scoring rules and a new scoring rule, the root score, are compared to maximum likelihood estimation. The results confirmed that for spatial Markov models the computation time for LOOS was much smaller than for maximum likelihood estimation. Furthermore, the standard deviations of parameter estimates were smaller for maximum likelihood estimation, although the differences often were small. The simulations also confirmed that the usage of a robust scoring rule results in robust LOOS estimates and that the robustness provides better predictive quality for spatial data with outliers. Finally, the new inference method was applied to ERA5 temperature reanalysis data for the contiguous United States and the average July temperature for the years 1940 to 2023, and this showed that the LOOS estimator provided parameter estimates that were more than a hundred times faster to compute compared to maximum-likelihood estimation, and resulted in a model with better predictive performance.

stat.ME↗

Locally tail-scale invariant scoring rules for evaluation of extreme value forecasts

Statistical analysis of extremes can be used to predict the probability of future extreme events, such as large rainfalls or devastating windstorms. The quality of these forecasts can be measured through scoring rules. Locally scale invariant scoring rules give equal importance to the forecasts at different locations regardless of differences in the prediction uncertainty. This is a useful feature when computing average scores but can be an unnecessarily strict requirement when mostly concerned with extremes. We propose the concept of local weight-scale invariance, describing scoring rules fulfilling local scale invariance in a certain region of interest, and as a special case local tail-scale invariance, for large events. Moreover, a new version of the weighted Continuous Ranked Probability score (wCRPS) called the scaled wCRPS (swCRPS) that possesses this property is developed and studied. The score is a suitable alternative for scoring extreme value models over areas with varying scale of extreme events, and we derive explicit formulas of the score for the Generalised Extreme Value distribution. The scoring rules are compared through simulation, and their usage is illustrated in modelling of extreme water levels, annual maximum rainfalls, and in an application to non-extreme forecast for the prediction of air pollution.

stat.ME↗

Real-time prediction of severe influenza epidemics using Extreme Value Statistics

Each year, seasonal influenza epidemics cause hundreds of thousands of deaths worldwide and put high loads on health care systems. A main concern for resource planning is the risk of exceptionally severe epidemics. Taking advantage of recent results on multivariate Generalized Pareto models in Extreme Value Statistics we develop methods for real-time prediction of the risk that an ongoing influenza epidemic will be exceptionally severe and for real-time detection of anomalous epidemics and use them for prediction and detection of anomalies for influenza epidemics in France. Quality of predictions is assessed on observed and simulated data.

stat.AP↗

Is there a cap on longevity? A statistical review

There is sustained and widespread interest in understanding the limit, if any, to the human lifespan. Apart from its intrinsic and biological interest, changes in survival in old age have implications for the sustainability of social security systems. A central question is whether the endpoint of the underlying lifetime distribution is finite. Recent analyses of data on the oldest human lifetimes have led to competing claims about survival and to some controversy, due in part to incorrect statistical analysis. This paper discusses the particularities of such data, outlines correct ways of handling them and presents suitable models and methods for their analysis. We provide a critical assessment of some earlier work and illustrate the ideas through reanalysis of semi-supercentenarian lifetime data. Our analysis suggests that remaining life-length after age 109 is exponentially distributed, and that any upper limit lies well beyond the highest lifetime yet reliably recorded. Lower limits to 95% confidence intervals for the human lifespan are around 130 years, and point estimates typically indicate no upper limit at all.

stat.AP↗

Human mortality at extreme age

We use a combination of extreme value theory, survival analysis and computer-intensive methods to analyze the mortality of Italian and French semi-supercentenarians for whom there are validated records. After accounting for the effects of the sampling frame, there appears to be a constant rate of mortality beyond age 108 years and no difference between countries and cohorts. These findings are consistent with previous work based on the International Database on Longevity and suggest that any physical upper bound for humans is so large that it is unlikely to be approached. There is no evidence of differences in survival between women and men after age 108 in the Italian data and the International Database on Longevity; however survival is lower for men in the French data.

stat.AP↗

Limit theorems for empirical processes of cluster functionals

Let $(X_{n,i})_{1\le i\le n,n\in\mathbb{N}}$ be a triangular array of row-wise stationary $\mathbb{R}^d$-valued random variables. We use a "blocks method" to define clusters of extreme values: the rows of $(X_{n,i})$ are divided into $m_n$ blocks $(Y_{n,j})$, and if a block contains at least one extreme value, the block is considered to contain a cluster. The cluster starts at the first extreme value in the block and ends at the last one. The main results are uniform central limit theorems for empirical processes $Z_n(f):=\frac{1}{\sqrt {nv_n}}\sum_{j=1}^{m_n}(f(Y_{n,j})-Ef(Y_{n,j})),$ for $v_n=P\{X_{n,i}\neq0\}$ and $f$ belonging to classes of cluster functionals, that is, functions of the blocks $Y_{n,j}$ which only depend on the cluster values and which are equal to 0 if $Y_{n,j}$ does not contain a cluster. Conditions for finite-dimensional convergence include $β$-mixing, suitable Lindeberg conditions and convergence of covariances. To obtain full uniform convergence, we use either "bracketing entropy" or bounds on covering numbers with respect to a random semi-metric. The latter makes it possible to bring the powerful Vapnik--Červonenkis theory to bear. Applications include multivariate tail empirical processes and empirical processes of cluster values and of order statistics in clusters. Although our main field of applications is the analysis of extreme values, the theory can be applied more generally to rare events occurring, for example, in nonparametric curve estimation.

math.ST↗

Peaks over thresholds modelling with multivariate generalized Pareto distributions

When assessing the impact of extreme events, it is often not just a single component, but the combined behaviour of several components which is important. Statistical modelling using multivariate generalized Pareto (GP) distributions constitutes the multivariate analogue of univariate peaks over thresholds modelling, which is widely used in finance and engineering. We develop general methods for construction of multivariate GP distributions and use them to create a variety of new statistical models. A censored likelihood procedure is proposed to make inference on these models, together with a threshold selection procedure, goodness-of-fit diagnostics, and a computationally tractable strategy for model selection. The models are fitted to returns of stock prices of four UK-based banks and to rainfall data in the context of landslide risk estimation. Supplementary materials and codes are available online.

stat.ME↗

Human life is unlimited - but short

Does the human lifespan have an impenetrable biological upper limit which ultimately will stop further increase in life lengths? This question is important for understanding aging, and for society, and has led to intense controversies. Demographic data for humans has been interpreted as showing existence of a limit, or even as an indication of a decreasing limit, but also as evidence that a limit does not exist. This paper studies what can be inferred from data about human mortality at extreme age. We show that in western countries and Japan and after age 110 the probability of dying is about 47% per year. Hence there is no finite upper limit to the human lifespan. Still, given the present stage of biotechnology, it is unlikely that during the next 25 years anyone will live longer than 128 years in these countries. Data, remarkably, shows no difference in mortality after age 110 between sexes, between ages, or between different lifestyles or genetic backgrounds. These results, and the analysis methods developed in this paper, can help testing biological theories of ageing and aid confirmation of success of efforts to find a cure for ageing.

stat.AP↗

Multivariate generalized Pareto distributions: parametrizations, representations, and properties

Multivariate generalized Pareto distributions arise as the limit distributions of exceedances over multivariate thresholds of random vectors in the domain of attraction of a max-stable distribution. These distributions can be parametrized and represented in a number of different ways. Moreover, generalized Pareto distributions enjoy a number of interesting stability properties. An overview of the main features of such distributions are given, expressed compactly in several parametrizations, giving the potential user of these distributions a convenient catalogue of ways to handle and work with generalized Pareto distributions.

math.ST↗

Multivariate peaks over thresholds models

Multivariate peaks over thresholds modeling based on generalized Pareto distributions has up to now only been used in few and mostly 2-dimensional situations. This paper contributes theoretical understanding, physically based models, inference tools, and simulation methods to support routine use, with an aim at higher dimensions. We derive a general point process model for extreme episodes in data, and show how conditioning the distribution of extreme episodes on threshold exceedance gives four basic representations of the family of generalized Pareto distributions. The first representation is constructed on the real scale of the observations. The second one starts with a model on a standard exponential scale which then is transformed to the real scale. The third and fourth are reformulations of a spectral representation proposed in A. Ferreira and L. de Haan [Bernoulli 20 (2014) 1717--1737]. Numerically tractable forms of densities and censored densities are found and give tools for flexible parametric likelihood inference. New simulation algorithms, explicit formulas for probabilities and conditional probabilities, and conditions which make the conditional distribution of weighted component sums generalized Pareto are derived.

math.PR↗

Error distributions for random grid approximations of multidimensional stochastic integrals

This paper proves joint convergence of the approximation error for several stochastic integrals with respect to local Brownian semimartingales, for nonequidistant and random grids. The conditions needed for convergence are that the Lebesgue integrals of the integrands tend uniformly to zero and that the squared variation and covariation processes converge. The paper also provides tools which simplify checking these conditions and which extend the range for the results. These results are used to prove an explicit limit theorem for random grid approximations of integrals based on solutions of multidimensional SDEs, and to find ways to "design" and optimize the distribution of the approximation error. As examples we briefly discuss strategies for discrete option hedging.

math.PR↗

Models for dependent extremes using stable mixtures

This paper unifies and extends results on a class of multivariate Extreme Value (EV) models studied by Hougaard, Crowder, and Tawn. In these models both unconditional and conditional distributions are EV, and all lower-dimensional marginals and maxima belong to the class. This leads to substantial economies of understanding, analysis and prediction. One interpretation of the models is as size mixtures of EV distributions, where the mixing is by positive stable distributions. A second interpretation is as exponential-stable location mixtures (for Gumbel) or as power-stable scale mixtures (for non-Gumbel EV distributions). A third interpretation is through a Peaks over Thresholds model with a positive stable intensity. The mixing variables are used as a modeling tool and for better understanding and model checking. We study extreme value analogues of components of variance models, and new time series, spatial, and continuous parameter models for extreme values. The results are applied to data from a pitting corrosion investigation.

stat.ME↗