SearcharxivSearch

arXiv subjects

Manuel Hentschel

Publications and source records attributed to Manuel Hentschel.

6 recordsLinked to original sources

An Explicit Link between Extreme Value Theory and Compositional Data Analysis

Extreme value theory and compositional data analysis both study settings where relative information plays a central role. In multivariate extreme value theory, threshold exceedance limits satisfy homogeneity properties that separate the radial size of an extreme event from its relative profile. In compositional data analysis, positive vectors are analysed up to multiplicative scale, and inference is based on ratios or log-ratios between components. Consequently, both fields have developed several covariance and dependence representations of the underlying relative structure. In the Hüsler-Reiss model for extremes, these include variogram, covariance, and precision parametrizations. In compositional data analysis, analogous representations arise from pairwise log-ratios, centred log-ratios, and additive log-ratios. We establish an explicit link between the two fields that relates these different representations by a small set of simple transformations, including oblique projections, Hüsler-Reiss inverses, and the variogram map. From a methodological perspective, leveraging this algebraic connection enables the transfer of statistical approaches from one field to the other. For instance, we introduce intrinsic logistic-normal graphical models for compositional data, which are based on Hüsler-Reiss graphical models for extremes. Conversely, we explore how dimensionality reduction methods from compositional data analysis can be applied to the analysis of multivariate extremes.

stat.ME

Directional variograms for multivariate extremes

Multivariate generalized Pareto distributions arise as limits of threshold exceedances and form a central model class for multivariate extremes. Existing inference methods based on the extremal variogram condition on the value of a single component, which can be statistically suboptimal. We generalize this approach by conditioning the multivariate generalized Pareto random vector $Y$ to lie on arbitrary half-spaces. Specifically, for a direction vector $v$, we introduce the random vector $Y^v = (Y \mid v^\top Y > 0)$ and define the associated $v$-variogram $Γ_{ij}^v=\mathrm{Var}(Y_i^v-Y_j^v)$. We establish the decomposition $Y^v \stackrel{d}{=} W^v+E\mathbf{1}$ into the so-called $v$-extremal function $W^v$ and an independent exponential random variable $E$, and derive several results relating these random variables to each other. For logistic, Dirichlet, and Hüsler-Reiss multivariate generalized Pareto models, we derive closed-form expressions for $Γ^v$. In the Hüsler-Reiss case, we further derive new density representations and identify a distinguished resistance-curvature vector $v_0$ that uniquely centers the Gaussian law of $W^{v_0}$ while characterizing the least-mass half-space. On the statistical side, we introduce empirical $v$-variograms and show in a simulation study that the choice of $v$ induces a pronounced bias-variance trade-off that is strongly related to the mass of the conditioning half-space. Moreover, combining information across multiple directions $v$ can substantially reduce estimation variance relative to methods based on a single vector.

stat.ME

Theoretical guarantees for neural estimators in parametric statistics

Neural estimators are simulation-based estimators for the parameters of a family of statistical models, which build a direct mapping from the sample to the parameter vector. They benefit from the versatility of available network architectures and efficient training methods developed in the field of deep learning. Neural estimators are amortized in the sense that, once trained, they can be applied to any new data set with almost no computational cost. While many papers have shown very good performance of these methods in simulation studies and real-world applications, so far no statistical guarantees are available to support these observations theoretically. In this work, we study the risk of neural estimators by decomposing it into several terms that can be analyzed separately. We formulate easy-to-check assumptions ensuring that each term converges to zero, and we verify them for popular applications of neural estimators. Our results provide a general recipe to derive theoretical guarantees also for broader classes of architectures and estimation problems.

stat.ML

Modeling Extreme Events: Univariate and Multivariate Data-Driven Approaches

This article summarizes the contribution of team genEVA to the EVA (2023) Conference Data Challenge. The challenge comprises four individual tasks, with two focused on univariate extremes and two related to multivariate extremes. In the first univariate assignment, we estimate a conditional extremal quantile using a quantile regression approach with neural networks. For the second, we develop a fine-tuning procedure for improved extremal quantile estimation with a given conservative loss function. In the first multivariate sub-challenge, we approximate the data-generating process with a copula model. In the remaining task, we use clustering to separate a high-dimensional problem into approximately independent components. Overall, competitive results were achieved for all challenges, and our approaches for the univariate tasks yielded the most accurate quantile estimates in the competition.

stat.ME

Graphical models for multivariate extremes

Graphical models in extremes have emerged as a diverse and quickly expanding research area in extremal dependence modeling. They allow for parsimonious statistical methodology and are particularly suited for enforcing sparsity in high-dimensional problems. In this work, we provide the fundamental concepts of extremal graphical models and discuss recent advances in the field. Different existing perspectives on graphical extremes are presented in a unified way through graphical models for exponent measures. We discuss the important cases of nonparametric extremal graphical models on simple graph structures, and the parametric class of Hüsler--Reiss models on arbitrary undirected graphs. In both cases, we describe model properties, methods for statistical inference on known graph structures, and structure learning algorithms when the graph is unknown. We illustrate different methods in an application to flight delay data at US airports.

stat.ME

Statistical Inference for Hüsler-Reiss Graphical Models Through Matrix Completions

The severity of multivariate extreme events is driven by the dependence between the largest marginal observations. The Hüsler-Reiss distribution is a versatile model for this extremal dependence, and it is usually parameterized by a variogram matrix. In order to represent conditional independence relations and obtain sparse parameterizations, we introduce the novel Hüsler-Reiss precision matrix. Similarly to the Gaussian case, this matrix appears naturally in density representations of the Hüsler-Reiss Pareto distribution and encodes the extremal graphical structure through its zero pattern. For a given, arbitrary graph we prove the existence and uniqueness of the completion of a partially specified Hüsler-Reiss variogram matrix so that its precision matrix has zeros on non-edges in the graph. Using suitable estimators for the parameters on the edges, our theory provides the first consistent estimator of graph structured Hüsler-Reiss distributions. If the graph is unknown, our method can be combined with recent structure learning algorithms to jointly infer the graph and the corresponding parameter matrix. Based on our methodology, we propose new tools for statistical inference of sparse Hüsler-Reiss models and illustrate them on large flight delay data in the U.S., as well as Danube river flow data.

stat.ME