SearcharxivSearch

arXiv subjects

Maud Thomas

Publications and source records attributed to Maud Thomas.

9 recordsLinked to original sources

Assessing Extreme Risk using Stochastic Simulation of Extremes

Risk management is particularly concerned with extreme events, but analysing these events is often hindered by the scarcity of data, especially in a multivariate context. This data scarcity complicates risk management efforts. Various tools can assess the risk posed by extreme events, even under extraordinary circumstances. This paper studies the evaluation of univariate risk for a given risk factor using metrics that account for its asymptotic dependence on other risk factors. Data availability is crucial, particularly for extreme events where it is often limited by the nature of the phenomenon itself, making estimation challenging. To address this issue, two non-parametric simulation algorithms based on multivariate extreme theory are developed. These algorithms aim to extend a sample of extremes jointly and conditionally for asymptotically dependent variables using stochastic simulation and multivariate Generalised Pareto Distributions. The approach is illustrated with numerical analyses of both simulated and real data to assess the accuracy of extreme risk metric estimations.

stat.ME

Tree-based conditional copula estimation

This paper proposes a regression tree procedure to estimate conditional copulas. The associated algorithm determines classes of observations based on covariate values and fits a simple parametric copula model on each class. The association parameter changes from one class to another, allowing for non-linearity in the dependence structure modeling. It also allows the definition of classes of observations on which the so-called "simplifying assumption" [see Derumigny and Fermanian, 2017] holds reasonably well. When considering observations belonging to a given class separately, the association parameter no longer depends on the covariates according to our model. In this paper, we derive asymptotic consistency results for the regression tree procedure and show that the proposed pruning methodology, that is the model selection techniques selecting the appropriate number of classes, is optimal in some sense. Simulations provide finite sample results and an analysis of data of cases of human influenza presents the practical behavior of the procedure.

math.ST

Robust and non asymptotic estimation of probability weighted moments with application to extreme value analysis

In extreme value theory and other related risk analysis fields, probability weighted moments (PWM) have been frequently used to estimate the parameters of classical extreme value distributions. This method-of-moment technique can be applied when second moments are finite, a reasonable assumption in many environmental domains like climatological and hydrological studies. Three advantages of PWM estimators can be put forward: their simple interpretations, their rapid numerical implementation and their close connection to the well-studied class of U-statistics. Concerning the later, this connection leads to precise asymptotic properties, but non asymptotic bounds have been lacking when off-the-shelf techniques (Chernoff method) cannot be applied, as exponential moment assumptions become unrealistic in many extreme value settings. In addition, large values analysis is not immune to the undesirable effect of outliers, for example, defective readings in satellite measurements or possible anomalies in climate model runs. Recently, the treatment of outliers has sparked some interest in extreme value theory, but results about finite sample bounds in a robust extreme value theory context are yet to be found, in particular for PWMs or tail index estimators. In this work, we propose a new class of robust PWM estimators, inspired by the median-of-means framework of Devroye et al. (2016). This class of robust estimators is shown to satisfy a sub-Gaussian inequality when the assumption of finite second moments holds. Such non asymptotic bounds are also derived under the general contamination model. Our main proposition confirms theoretically a trade-off between efficiency and robustness. Our simulation study indicates that, while classical estimators of PWMs can be highly sensitive to outliers.

math.ST

Parametric insurance for extreme risks: the challenge of properly covering severe claims

Parametric insurance has emerged as a practical way to cover risks that may be difficult to assess. By introducing a parameter that triggers compensation and allows the insurer to determine a payment without estimating the actual loss, these products simplify the compensation process, and provide easily traceable indicators to perform risk management. On the other hand, this parameter may sometimes deviate from its intended purpose, and may not always accurately represent the basic risk. In this paper, we provide theoretical results that investigate the behavior of parametric insurance products when faced with large claims. In particular, these results measure the difference between the actual loss and the parameter in a generic situation, with a particular focus on heavy-tailed losses. These results may help to anticipate, in presence of heavy-tail phenomena, how parametric products should be supplemented by additional compensation mechanisms in case of large claims. Simulation studies, that complement the analysis, show the importance of nonlinear dependence measures in providing a good protection over the whole distribution.

stat.AP

Anomaly Detection on Financial Time Series by Principal Component Analysis and Neural Networks

A major concern when dealing with financial time series involving a wide variety ofmarket risk factors is the presence of anomalies. These induce a miscalibration of the models used toquantify and manage risk, resulting in potential erroneous risk measures. We propose an approachthat aims to improve anomaly detection in financial time series, overcoming most of the inherentdifficulties. Valuable features are extracted from the time series by compressing and reconstructingthe data through principal component analysis. We then define an anomaly score using a feedforwardneural network. A time series is considered to be contaminated when its anomaly score exceeds agiven cutoff value. This cutoff value is not a hand-set parameter but rather is calibrated as a neuralnetwork parameter throughout the minimization of a customized loss function. The efficiency of theproposed approach compared to several well-known anomaly detection algorithms is numericallydemonstrated on both synthetic and real data sets, with high and stable performance being achievedwith the PCA NN approach. We show that value-at-risk estimation errors are reduced when theproposed anomaly detection model is used with a basic imputation approach to correct the anomaly.

q-fin.ST

Generalized Pareto Regression Trees for extreme events analysis

In this paper, we provide finite sample results to assess the consistency of Generalized Pareto regression trees, as tools to perform extreme value regression. The results that we provide are obtained from concentration inequalities, and are valid for a finite sample size, taking into account a misspecification bias that arises from the use of a "Peaks over Threshold" approach. The properties that we derive also legitimate the pruning strategies (i.e. the model selection rules) used to select a proper tree that achieves compromise between bias and variance. The methodology is illustrated through a simulation study, and a real data application in insurance against natural disasters.

math.ST

Real-time prediction of severe influenza epidemics using Extreme Value Statistics

Each year, seasonal influenza epidemics cause hundreds of thousands of deaths worldwide and put high loads on health care systems. A main concern for resource planning is the risk of exceptionally severe epidemics. Taking advantage of recent results on multivariate Generalized Pareto models in Extreme Value Statistics we develop methods for real-time prediction of the risk that an ongoing influenza epidemic will be exceptionally severe and for real-time detection of anomalous epidemics and use them for prediction and detection of anomalies for influenza epidemics in France. Quality of predictions is assessed on observed and simulated data.

stat.AP

Tail index estimation, concentration and adaptivity

This paper presents an adaptive version of the Hill estimator based on Lespki's model selection method. This simple data-driven index selection method is shown to satisfy an oracle inequality and is checked to achieve the lower bound recently derived by Carpentier and Kim. In order to establish the oracle inequality, we derive non-asymptotic variance bounds and concentration inequalities for Hill estimators. These concentration inequalities are derived from Talagrand's concentration inequality for smooth functions of independent exponentially distributed random variables combined with three tools of Extreme Value Theory: the quantile transform, Karamata's representation of slowly varying functions, and R\'enyi's characterisation of the order statistics of exponential samples. The performance of this computationally and conceptually simple method is illustrated using Monte-Carlo simulations.

math.ST

Concentration inequalities for order statistics

This note describes non-asymptotic variance and tail bounds for order statistics of samples of independent identically distributed random variables. Those bounds are checked to be asymptotically tight when the sampling distribution belongs to a maximum domain of attraction. If the sampling distribution has non-decreasing hazard rate (this includes the Gaussian distribution), we derive an exponential Efron-Stein inequality for order statistics: an inequality connecting the logarithmic moment generating function of centered order statistics with exponential moments of Efron-Stein (jackknife) estimates of variance. We use this general connection to derive variance and tail bounds for order statistics of Gaussian sample. Those bounds are not within the scope of the Tsirelson-Ibragimov-Sudakov Gaussian concentration inequality. Proofs are elementary and combine Rényi's representation of order statistics and the so-called entropy approach to concentration inequalities popularized by M. Ledoux.

math.PR