SearcharxivSearch

arXiv subjects

Holger Drees

Publications and source records attributed to Holger Drees.

At least 19 recordsLinked to original sources

Asymptotic Behavior of Principal Component Projections for Multivariate Extremes

The extremal dependence structure of a regularly varying $d$-dimensional random vector can be described by its angular measure. The standard nonparametric estimator of this measure is the empirical measure of the observed angles of the $k$ random vectors with largest norm, for a suitably chosen number $k$. Due to the curse of dimensionality, for moderate or large $d$, this estimator is often inaccurate. If the angular measure is concentrated on a vicinity of a lower dimensional subspace, then first projecting the data on a lower dimensional subspace obtained by a principal component analysis of the angles of extreme observations can substantially improve the performance of the estimator. We derive the asymptotic behavior of such PCA projections and the resulting excess risk. In particular, it is shown that, under mild conditions, the excess risk (as a function of $k$) decreases much faster than it was suggested by empirical risk bounds obtained in \cite{DS21}. Moreover, functional limit theorems for local empirical processes of the (empirical) reconstruction error of projections uniformly over neighborhoods of the true optimal projection are established. Based on these asymptotic results, we propose a data-driven method to select the dimension of the projection space. Finally, the finite sample performance of resulting estimators is examined in a simulation study.

math.ST

Statistical Inference on a Changing Extremal Dependence Structure

We analyze the extreme value dependence of independent, not necessarily identically distributed multivariate regularly varying random vectors. More specifically, we propose estimators of the spectral measure locally at some time point and of the spectral measures integrated over time. The uniform asymptotic normality of these estimators is proved under suitable nonparametric smoothness and regularity assumptions. We then use the process convergence of the integrated spectral measure to devise consistent tests for the null hypothesis that the spectral measure does not change over time. The finite sample performance of these tests is investigated in Monte Carlo simulations.

math.ST

Cluster based inference for extremes of time series

We introduce a new type of estimator for the spectral tail process of a regularly varying time series. The approach is based on a characterizing invariance property of the spectral tail process, which is incorporated into the new estimator via a projection technique. We show uniform asymptotic normality of this estimator, both in the case of known and of unknown index of regular variation. In a simulation study the new procedure shows a more stable performance than previously proposed estimators.

math.ST

Asymptotics for sliding blocks estimators of rare events

Drees and Rootz\'en (2010) have established limit theorems for a general class of empirical processes of statistics that are useful for the extreme value analysis of time series, but do not apply to statistics of sliding blocks, including so-called runs estimators. We generalize these results to empirical processes which cover both the class considered by Drees and Rootz\'en (2010) and processes of sliding blocks statistics. Using this approach, one can analyze different types of statistics in a unified framework. We show that statistics based on sliding blocks are asymptotically normal with an asymptotic variance which, under rather mild conditions, is smaller than or equal to the asymptotic variance of the corresponding estimator based on disjoint blocks. Finally, the general theory is applied to three well-known estimators of the extremal index. It turns out that they all have the same limit distribution, a fact which has so far been overlooked in the literature.

math.ST

Principal Component Analysis for Multivariate Extremes

The first order behavior of multivariate heavy-tailed random vectors above large radial thresholds is ruled by a limit measure in a regular variation framework. For a high dimensional vector, a reasonable assumption is that the support of this measure is concentrated on a lower dimensional subspace, meaning that certain linear combinations of the components are much likelier to be large than others. Identifying this subspace and thus reducing the dimension will facilitate a refined statistical analysis. In this work we apply Principal Component Analysis (PCA) to a re-scaled version of radially thresholded observations. Within the statistical learning framework of empirical risk minimization, our main focus is to analyze the squared reconstruction error for the exceedances over large radial thresholds. We prove that the empirical risk converges to the true risk, uniformly over all projection subspaces. As a consequence, the best projection subspace is shown to converge in probability to the optimal one, in terms of the Hausdorff distance between their intersections with the unit sphere. In addition, if the exceedances are re-scaled to the unit ball, we obtain finite sample uniform guarantees to the reconstruction error pertaining to the estimated projection sub-space. Numerical experiments illustrate the relevance of the proposed framework for practical purposes.

math.ST

Peak-over-Threshold Estimators for Spectral Tail Processes: Random vs Deterministic Thresholds

The extreme value dependence of regularly varying stationary time series can be described by the spectral tail process. Drees, Segers and Warchol [Extremes 18(3): 369--402, 2015] proposed estimators of the marginal distributions of this process based on exceedances over high deterministic thresholds and analyzed their asymptotic behavior. In practice, however, versions of the estimators are applied which use exceedances over random thresholds like intermediate order statistics. We prove that these modified estimators have the same limit distributions. This finding is corroborated in a simulation study, but the version using order statistics performs a bit better for finite samples.

math.ST

On a minimum distance procedure for threshold selection in tail analysis

Power-law distributions have been widely observed in different areas of scientific research. Practical estimation issues include how to select a threshold above which observations follow a power-law distribution and then how to estimate the power-law tail index. A minimum distance selection procedure (MDSP) is proposed in Clauset et al. (2009) and has been widely adopted in practice, especially in the analyses of social networks. However, theoretical justifications for this selection procedure remain scant. In this paper, we study the asymptotic behavior of the selected threshold and the corresponding power-law index given by the MDSP. We find that the MDSP tends to choose too high a threshold level and leads to Hill estimates with large variances and root mean squared errors for simulated data with Pareto-like tails.

math.ST

Atomic-scale imaging of the surface dipole distribution of stepped surfaces

Stepped well-ordered semiconductor surfaces are important as nanotemplates for the fabrication of one-dimensional nanostructures which are candidates of intriguing electronic properties. Therefore a detailed understanding of the underlying stepped substrates is crucial for advances in this field. Although measurements of step edges are challenging for scanning force microscopy (SFM), here we present for the first time simultaneous atomically resolved SFM and Kelvin probe force microscopy (KPFM) images of a silicon vicinal surface. We find that the local contact potential difference is not homogeneous over all silicon atoms, contrary to the common understanding of the work function. For the interpretation of the data we performed density functional theory (DFT) calculations. We explain the atomic-scale electronic features by differences in the surface dipole distribution caused by a Smoluchowski-type effect. E.g., at step edges the partial negative charge is larger at the lower-lying atoms closer to the bulk than at the atoms more protruding towards the vacuum. This is the first manifestation of such type of effect on a semiconductor surface. The DFT images accurately reproduce the experiments even without including the tip in the calculations. This underlines that the high-resolution KPFM images indeed show the intrinsic properties of the surface and not only tip-surface interactions.

cond-mat.mes-hall

Atomically resolved scanning force studies of vicinal Si(111)

Well-ordered stepped semiconductor surfaces attract intense attention owing to the regular arrangements of their atomic steps that makes them perfect templates for the growth of one- dimensional systems, e.g. nanowires. Here, we report on the atomic structure of the vicinal Si(111) surface with 10 degree miscut investigated by a joint frequency-modulation scanning force microscopy (FM-SFM) and ab initio approach. This popular stepped surface contains 7 x 7-reconstructed terraces oriented along the Si(111) direction, separated by a stepped region. Recently, the atomic structure of this triple step based on scanning tunneling microscopy (STM) images has been subject of debate. Unlike STM, SFM atomic resolution capability arises from chemical bonding of the tip apex with the surface atoms. Thus, for surfaces with a corrugated density of states such as semiconductors, SFM provides complementary information to STM and partially removes the dependency of the topography on the electronic structure. Our FM-SFM images with unprecedented spatial resolution on steps confirm the model based on a (7 7 10) orientation of the surface and reveal structural details of this surface. Two different FM-SFM contrasts together with density functional theory calculations explain the presence of defects, buckling and filling asymmetries on the surface. Our results evidence the important role of charge transfers between adatoms, restatoms, and dimers in the stabilisation of the structure of the vicinal surface.

cond-mat.mes-hall

Extreme Value Estimation for Discretely Sampled Continuous Processes

In environmental applications of extreme value statistics, the underlying stochastic process is often modeled either as a max-stable process in continuous time/space or as a process in the domain of attraction of such a max-stable process. In practice, however, the processes are typically only observed at discrete points and one has to resort to interpolation to fill in the gaps. We discuss the influence of such an interpolation on estimators of marginal parameters as well as estimators of the exponent measure. In particular, natural conditions on the fineness of the observational scheme are developed which ensure that asymptotically the interpolated estimators behave in the same way as the estimators which use fully observed continuous processes.

math.ST

Conditional Extreme Value Models: Fallacies and Pitfalls

Conditional extreme value models have been introduced by Heffernan and Resnick (2007) to describe the asymptotic behavior of a random vector as one specific component becomes extreme. Obviously, this class of models is related to classical multivariate extreme value theory which describes the behavior of a random vector as its norm (and therefore at least one of its components) becomes extreme. However, it turns out that this relationship is rather subtle and sometimes contrary to intuition. We clarify the differences between the two approaches with the help of several illuminative (counter)examples. Furthermore, we discuss marginal standardization, which is a useful tool in classical multivariate extreme value theory but, as we point out, much less straightforward and sometimes even obscuring in conditional extreme value models. Finally, we indicate how, in some situations, a more comprehensive characterization of the asymptotic behavior can be obtained if the conditions of conditional extreme value models are relaxed so that the limit is no longer unique.

math.PR

Bootstrapping Empirical Processes of Cluster Functionals with Application to Extremograms

In the extreme value analysis of time series, not only the tail behavior is of interest, but also the serial dependence plays a crucial role. Drees and Rootz\'en (2010) established limit theorems for a general class of empirical processes of so-called cluster functionals which can be used to analyse various aspects of the extreme value behavior of mixing time series. However, usually the limit distribution is too complex to enable a direct construction of confidence regions. Therefore, we suggest a multiplier block bootstrap analog to the empirical processes of cluster functionals. It is shown that under virtually the same conditions as used by Drees and Rootz\'en (2010), conditionally on the data, the bootstrap processes converge to the same limit distribution. These general results are applied to construct confidence regions for the empirical extremogram introduced by Davis and Mikosch (2009). In a simulation study, the confidence intervals constructed by our multiplier block bootstrap approach compare favorably to the stationary bootstrap proposed by Davis et al.\ (2012).

math.ST

Joint exceedances of random products

We analyze the joint extremal behavior of $n$ random products of the form $\prod_{j=1}^m X_j^{a_{ij}}, 1 \leq i \leq n,$ for non-negative, independent regularly varying random variables $X_1, \ldots, X_m$ and general coefficients $a_{ij} \in \mathbb{R}$. Products of this form appear for example if one observes a linear time series with gamma type innovations at $n$ points in time. We combine arguments of linear optimization and a generalized concept of regular variation on cones to show that the asymptotic behavior of joint exceedance probabilities of these products is determined by the solution of a linear program related to the matrix $\mathbf{A}=(a_{ij})$.

math.PR

Statistics for Tail Processes of Markov Chains

At high levels, the asymptotic distribution of a stationary, regularly varying Markov chain is conveniently given by its tail process. The latter takes the form of a geometric random walk, the increment distribution depending on the sign of the process at the current state and on the flow of time, either forward or backward. Estimation of the tail process provides a nonparametric approach to analyze extreme values. A duality between the distributions of the forward and backward increments provides additional information that can be exploited in the construction of more efficient estimators. The large-sample distribution of such estimators is derived via empirical process theory for cluster functionals. Their finite-sample performance is evaluated via Monte Carlo simulations involving copula-based Markov models and solutions to stochastic recurrence equations. The estimators are applied to stock price data to study the absence or presence of symmetries in the succession of large gains and losses.

stat.ME

Hypotheses tests in boundary regression models

Consider a nonparametric regression model with one-sided errors and regression function in a general H\"older class. We estimate the regression function via minimization of the local integral of a polynomial approximation. We show uniform rates of convergence for the simple regression estimator as well as for a smooth version. These rates carry over to mean regression models with a symmetric and bounded error distribution. In such a setting, one obtains faster rates for irregular error distributions concentrating sufficient mass near the endpoints than for the usual regular distributions. The results are applied to prove asymptotic $\sqrt{n}$-equivalence of a residual-based (sequential) empirical distribution function to the (sequential) empirical distribution function of unobserved errors in the case of irregular error distributions. This result is remarkably different from corresponding results in mean regression with regular errors. It can readily be applied to develop goodness-of-fit tests for the error distribution. We present some examples and investigate the small sample performance in a simulation study. We further discuss asymptotically distribution-free hypotheses tests for independence of the error distribution from the points of measurement and for monotonicity of the boundary function as well.

stat.ME

A stochastic volatility model with flexible extremal dependence structure

Stochastic volatility processes with heavy-tailed innovations are a well-known model for financial time series. In these models, the extremes of the log returns are mainly driven by the extremes of the i.i.d. innovation sequence which leads to a very strong form of asymptotic independence, that is, the coefficient of tail dependence is equal to $1/2$ for all positive lags. We propose an alternative class of stochastic volatility models with heavy-tailed volatilities and examine their extreme value behavior. In particular, it is shown that, while lagged extreme observations are typically asymptotically independent, their coefficient of tail dependence can take on any value between $1/2$ (corresponding to exact independence) and 1 (related to asymptotic dependence). Hence, this class allows for a much more flexible extremal dependence between consecutive observations than classical SV models and can thus describe the observed clustering of financial returns more realistically. The extremal dependence structure of lagged observations is analyzed in the framework of regular variation on the cone $(0,\infty)^d$. As two auxiliary results which are of interest on their own we derive a new Breiman-type theorem about regular variation on $(0,\infty)^d$ for products of a random matrix and a regularly varying random vector and a statement about the joint extremal behavior of products of i.i.d. regularly varying random variables.

math.PR

Extreme value analysis of actuarial risks: estimation and model validation

We give an overview of several aspects arising in the statistical analysis of extreme risks with actuarial applications in view. In particular it is demonstrated that empirical process theory is a very powerful tool, both for the asymptotic analysis of extreme value estimators and to devise tools for the validation of the underlying model assumptions. While the focus of the paper is on univariate tail risk analysis, the basic ideas of the analysis of the extremal dependence between different risks are also outlined. Here we emphasize some of the limitation of classical multivariate extreme value theory and sketch how a different model proposed by Ledford and Tawn can help to avoid pitfalls. Finally, these theoretical results are used to analyze a data set of large claim sizes from health insurance.

stat.ME