Searcharxiv⌕ Search

arXiv subjects

Johan Segers

Publications and source records attributed to Johan Segers.

At least 55 records · Page 3Linked to original sources

Tails of optimal transport plans for regularly varying probability measures

For the basic case of $L_2$ optimal transport between two probability measures on a Euclidean space, the regularity of the coupling measure and the transport map in the tail regions of these measures is studied. For this purpose, Robert McCann's classical existence and uniqueness results are extended to a class of possibly infinite measures, finite outside neighbourhoods of the origin. For convergent sequences of pairs of such measures, the stability of the multivalued transport maps is considered, and a useful notion of locally uniform convergence of these maps is verified under light assumptions. Applied to regularly varying probability measures, these general results imply the existence of tail limits of the transport plan and the coupling measure, these objects exhibiting distinct types of homogeneity.

math.PR↗

Bayesian model averaging over tree-based dependence structures for multivariate extremes

Describing the complex dependence structure of extreme phenomena is particularly challenging. To tackle this issue we develop a novel statistical algorithm that describes extremal dependence taking advantage of the inherent hierarchical dependence structure of the max-stable nested logistic distribution and that identifies possible clusters of extreme variables using reversible jump Markov chain Monte Carlo techniques. Parsimonious representations are achieved when clusters of extreme variables are found to be completely independent. Moreover, we significantly decrease the computational complexity of full likelihood inference by deriving a recursive formula for the nested logistic model likelihood. The algorithm performance is verified through extensive simulation experiments which also compare different likelihood procedures. The new methodology is used to investigate the dependence relationships between extreme concentration of multiple pollutants in California and how these pollutants are related to extreme weather conditions. Overall, we show that our approach allows for the representation of complex extremal dependence structures and has valid applications in multivariate data analysis, such as air pollution monitoring, where it can guide policymaking.

stat.ME↗

Identifying groups of variables with the potential of being large simultaneously

Identifying groups of variables that may be large simultaneously amounts to finding out which joint tail dependence coefficients of a multivariate distribution are positive. The asymptotic distribution of a vector of nonparametric, rank-based estimators of these coefficients justifies a stopping criterion in an algorithm that searches the collection of all possible groups of variables in a systematic way, from smaller groups to larger ones. The issue that the tolerance level in the stopping criterion should depend on the size of the groups is circumvented by the use of a conditional tail dependence coefficient. Alternatively, such stopping criteria can be based on limit distributions of rank-based estimators of the coefficient of tail dependence, quantifying the speed of decay of joint survival functions. Numerical experiments indicate that the algorithm's effectiveness for detecting tail-dependent groups of variables is highest when paired with a criterion based on a Hill-type estimator of the coefficient of tail dependence.

stat.ME↗

Inference for heavy tailed stationary time series based on sliding blocks

The block maxima method in extreme value theory consists of fitting an extreme value distribution to a sample of block maxima extracted from a time series. Traditionally, the maxima are taken over disjoint blocks of observations. Alternatively, the blocks can be chosen to slide through the observation period, yielding a larger number of overlapping blocks. Inference based on sliding blocks is found to be more efficient than inference based on disjoint blocks. The asymptotic variance of the maximum likelihood estimator of the Fréchet shape parameter is reduced by more than 18%. Interestingly, the amount of the efficiency gain is the same whatever the serial dependence of the underlying time series: as for disjoint blocks, the asymptotic distribution depends on the serial dependence only through the sequence of scaling constants. The findings are illustrated by simulation experiments and are applied to the estimation of high return levels of the daily log-returns of the Standard & Poor's 500 stock market index.

math.ST↗

Bayesian inference for bivariate ranks

A recommender system based on ranks is proposed, where an expert's ranking of a set of objects and a user's ranking of a subset of those objects are combined to make a prediction of the user's ranking of all objects. The rankings are assumed to be induced by latent continuous variables corresponding to the grades assigned by the expert and the user to the objects. The dependence between the expert and user grades is modelled by a copula in some parametric family. Given a prior distribution on the copula parameter, the user's complete ranking is predicted by the mode of the posterior predictive distribution of the user's complete ranking conditional on the expert's complete and the user's incomplete rankings. Various Markov chain Monte-Carlo algorithms are proposed to approximate the predictive distribution or only its mode. The predictive distribution can be obtained exactly for the Farlie-Gumbel-Morgenstern copula family, providing a benchmark for the approximation accuracy of the algorithms. The method is applied to the MovieLens 100k dataset with a Gaussian copula modelling dependence between the expert's and user's grades.

stat.ML↗

Peaks over thresholds modelling with multivariate generalized Pareto distributions

When assessing the impact of extreme events, it is often not just a single component, but the combined behaviour of several components which is important. Statistical modelling using multivariate generalized Pareto (GP) distributions constitutes the multivariate analogue of univariate peaks over thresholds modelling, which is widely used in finance and engineering. We develop general methods for construction of multivariate GP distributions and use them to create a variety of new statistical models. A censored likelihood procedure is proposed to make inference on these models, together with a threshold selection procedure, goodness-of-fit diagnostics, and a computationally tractable strategy for model selection. The models are fitted to returns of stock prices of four UK-based banks and to rainfall data in the context of landslide risk estimation. Supplementary materials and codes are available online.

stat.ME↗

Weak convergence of the weighted empirical beta copula process

The empirical copula has proved to be useful in the construction and understanding of many statistical procedures related to dependence within random vectors. The empirical beta copula is a smoothed version of the empirical copula that enjoys better finite-sample properties. At the core lie fundamental results on the weak convergence of the empirical copula and empirical beta copula processes. Their scope of application can be increased by considering weighted versions of these processes. In this paper we show weak convergence for the weighted empirical beta copula process. The weak convergence result for the weighted empirical beta copula process is stronger than the one for the empirical copula and its use is more straightforward. The simplicity of its application is illustrated for weighted Cramér--von Mises tests for independence and for the estimation of the Pickands dependence function of an extreme-value copula.

math.ST↗

On the longest gap between power-rate arrivals

Let $L_t$ be the longest gap before time $t$ in an inhomogeneous Poisson process with rate function $λ_t$ proportional to $t^{α-1}$ for some $α\in(0,1)$. It is shown that $λ_tL_t-b_t$ has a limiting Gumbel distribution for suitable constants $b_t$ and that the distance of this longest gap from $t$ is asymptotically of the form $(t/\log t)E$ for an exponential random variable $E$. The analysis is performed via weak convergence of related point processes. Subject to a weak technical condition, the results are extended to include a slowly varying term in $λ_t$.

math.PR↗

An estimator of the stable tail dependence function based on the empirical beta copula

The replacement of indicator functions by integrated beta kernels in the definition of the empirical stable tail dependence function is shown to produce a smoothed version of the latter estimator with the same asymptotic distribution but superior finite-sample performance. The link of the new estimator with the empirical beta copula enables a simple but effective resampling scheme.

stat.ME↗

Multivariate generalized Pareto distributions: parametrizations, representations, and properties

Multivariate generalized Pareto distributions arise as the limit distributions of exceedances over multivariate thresholds of random vectors in the domain of attraction of a max-stable distribution. These distributions can be parametrized and represented in a number of different ways. Moreover, generalized Pareto distributions enjoy a number of interesting stability properties. An overview of the main features of such distributions are given, expressed compactly in several parametrizations, giving the potential user of these distributions a convenient catalogue of ways to handle and work with generalized Pareto distributions.

math.ST↗

On the weak convergence of the empirical conditional copula under a simplifying assumption

When the copula of the conditional distribution of two random variables given a covariate does not depend on the value of the covariate, two conflicting intuitions arise about the best possible rate of convergence attainable by nonparametric estimators of that copula. In the end, any such estimator must be based on the marginal conditional distribution functions of the two dependent variables given the covariate, and the best possible rates for estimating such localized objects is slower than the parametric one. However, the invariance of the conditional copula given the value of the covariate suggests the possibility of parametric convergence rates. The more optimistic intuition is shown to be correct, confirming a conjecture supported by extensive Monte Carlo simulations by I. Hobaek Haff and J. Segers [Computational Statistics and Data Analysis 84:1--13, 2015] and improving upon the nonparametric rate obtained theoretically by I. Gijbels, M. Omelka and N. Veraverbeke [Scandinavian Journal of Statistics 2015, to appear]. The novelty of the proposed approach lies in the double smoothing procedure employed for the estimator of the marginal cumulative distribution functions. Under mild conditions on the bandwidth sequence, the estimator is shown to take values in a certain class of smooth functions, the class having sufficiently small entropy for empirical process arguments to work. The copula estimator itself is asymptotically undistinguishable from a kind of oracle empirical copula, making it appear as if the marginal conditional distribution functions were known.

math.ST↗

Multivariate peaks over thresholds models

Multivariate peaks over thresholds modeling based on generalized Pareto distributions has up to now only been used in few and mostly 2-dimensional situations. This paper contributes theoretical understanding, physically based models, inference tools, and simulation methods to support routine use, with an aim at higher dimensions. We derive a general point process model for extreme episodes in data, and show how conditioning the distribution of extreme episodes on threshold exceedance gives four basic representations of the family of generalized Pareto distributions. The first representation is constructed on the real scale of the observations. The second one starts with a model on a standard exponential scale which then is transformed to the real scale. The third and fourth are reformulations of a spectral representation proposed in A. Ferreira and L. de Haan [Bernoulli 20 (2014) 1717--1737]. Numerically tractable forms of densities and censored densities are found and give tools for flexible parametric likelihood inference. New simulation algorithms, explicit formulas for probabilities and conditional probabilities, and conditions which make the conditional distribution of weighted component sums generalized Pareto are derived.

math.PR↗

On the maximum likelihood estimator for the Generalized Extreme-Value distribution

The vanilla method in univariate extreme-value theory consists of fitting the three-parameter Generalized Extreme-Value (GEV) distribution to a sample of block maxima. Despite claims to the contrary, the asymptotic normality of the maximum likelihood estimator has never been established. In this paper, a formal proof is given using a general result on the maximum likelihood estimator for parametric families that are differentiable in quadratic mean but whose supports depend on the parameter. An interesting side result concerns the (lack of) differentiability in quadratic mean of the GEV family.

math.ST↗

Polar decomposition of regularly varying time series in star-shaped metric spaces

There exist two ways of defining regular variation of a time series in a star-shaped metric space: either by the distributions of finite stretches of the series or by viewing the whole series as a single random element in a sequence space. The two definitions are shown to be equivalent. The introduction of a norm-like function, called modulus, yields a polar decomposition similar to the one in Euclidean spaces. The angular component of the time series, called angular or spectral tail process, captures all aspects of extremal dependence. The stationarity of the underlying series induces a transformation formula of the spectral tail process under time shifts.

math.PR↗

Marginal standardization of upper semicontinuous processes. with application to max-stable processes

Extreme-value theory for random vectors and stochastic processes with continuous trajectories is usually formulated for random objects all of whose univariate marginal distributions are identical. In the spirit of Sklar's theorem from copula theory, such marginal standardization is carried out by the pointwise probability integral transform. Certain situations, however, call for stochastic models whose trajectories are not continuous but merely upper semicontinuous (usc). Unfortunately, the pointwise application of the probability integral transform to a usc process does in general not preserve the upper semicontinuity of the trajectories. In the present work, we give sufficient conditions for marginal standardization of usc processes to be possible, and we state a partial extension of Sklar's theorem for usc processes. We specialize the results to max-stable processes whose marginal distributions and normalizing sequences are allowed to vary with the coordinate.

math.PR↗

The Empirical Beta Copula

Given a sample from a multivariate distribution $F$, the uniform random variates generated independently and rearranged in the order specified by the componentwise ranks of the original sample look like a sample from the copula of $F$. This idea can be regarded as a variant on Baker's [J. Multivariate Anal. 99 (2008) 2312--2327] copula construction and leads to the definition of the empirical beta copula. The latter turns out to be a particular case of the empirical Bernstein copula, the degrees of all Bernstein polynomials being equal to the sample size. Necessary and sufficient conditions are given for a Bernstein polynomial to be a copula. These imply that the empirical beta copula is a genuine copula. Furthermore, the empirical process based on the empirical Bernstein copula is shown to be asymptotically the same as the ordinary empirical copula process under assumptions which are significantly weaker than those given in Janssen, Swanepoel and Veraverbeke [J. Stat. Plan. Infer. 142 (2012) 1189--1197]. A Monte Carlo simulation study shows that the empirical beta copula outperforms the empirical copula and the empirical checkerboard copula in terms of both bias and variance. Compared with the empirical Bernstein copula with the smoothing rate suggested by Janssen et al., its finite-sample performance is still significantly better in several cases, especially in terms of bias.

math.ST↗

Maximum likelihood estimation for the Fréchet distribution based on block maxima extracted from a time series

The block maxima method in extreme-value analysis proceeds by fitting an extreme-value distribution to a sample of block maxima extracted from an observed stretch of a time series. The method is usually validated under two simplifying assumptions: the block maxima should be distributed according to an extreme-value distribution and the sample of block maxima should be independent. Both assumptions are only approximately true. For general triangular arrays of block maxima attracted to the Fréchet distribution, consistency and asymptotic normality is established for the maximum likelihood estimator of the parameters of the limiting Fréchet distribution. The results are specialized to the setting of block maxima extracted from a strictly stationary time series. The case where the underlying random variables are independent and identically distributed is further worked out in detail. The results are illustrated by theoretical examples and Monte Carlo simulations.

math.ST↗

A continuous updating weighted least squares estimator of tail dependence in high dimensions

Likelihood-based procedures are a common way to estimate tail dependence parameters. They are not applicable, however, in non-differentiable models such as those arising from recent max-linear structural equation models. Moreover, they can be hard to compute in higher dimensions. An adaptive weighted least-squares procedure matching nonparametric estimates of the stable tail dependence function with the corresponding values of a parametrically specified proposal yields a novel minimum-distance estimator. The estimator is easy to calculate and applies to a wide range of sampling schemes and tail dependence models. In large samples, it is asymptotically normal with an explicit and estimable covariance matrix. The minimum distance obtained forms the basis of a goodness-of-fit statistic whose asymptotic distribution is chi-square. Extensive Monte Carlo simulations confirm the excellent finite-sample performance of the estimator and demonstrate that it is a strong competitor to currently available methods. The estimator is then applied to disentangle sources of tail dependence in European stock markets.

stat.ME↗