SearcharxivSearch

arXiv subjects

Jan Beirlant

Publications and source records attributed to Jan Beirlant.

At least 19 recordsLinked to original sources

Statistics of Extremes for the Insurance Industry

We provide a survey of how techniques developed for the modelling of extremes naturally matter in insurance, and how they need to and can be adapted for the insurance applications. Topics covered include truncation, tempering, censoring and regression techniques. The discussed techniques are illustrated on concrete data sets.

q-fin.RM

Non-parametric cure models through extreme-value tail estimation

In survival analysis, the estimation of the proportion of subjects who will never experience the event of interest, termed the cure rate, has received considerable attention recently. Its estimation can be a particularly difficult task when follow-up is not sufficient, that is when the censoring mechanism has a smaller support than the distribution of the target data. In the latter case, non-parametric estimators were recently proposed using extreme value methodology, assuming that the distribution of the susceptible population is in the Fr\'echet or Gumbel max-domains of attraction. In this paper, we take the extreme value techniques one step further, to jointly estimate the cure rate and the extreme value index, using probability plotting methodology, and in particular using the full information contained in the top order statistics. In other words, under sufficient or insufficient follow-up, we reconstruct the immune proportion. To this end, a Peaks-over-Threshold approach is proposed under the Gumbel max-domain assumption. Next, the approach is also transferred to more specific models such as Pareto, log-normal and Weibull tail models, allowing to recognize the most important tail characteristics of the susceptible population. We establish the asymptotic behavior of our estimators under regularization. Though simulation studies, our estimators are show to rival and often outperform established models, even when purely considering cure rate estimation. Finally, we provide an application of our method to Norwegian birth registry data.

math.ST

A new class of copula regression models for modelling multivariate heavy-tailed data

A new class of copulas, termed the MGL copula class, is introduced. The new copula originates from extracting the dependence function of the multivariate generalized log-Moyal-gamma distribution whose marginals follow the univariate generalized log-Moyal-gamma (GLMGA) distribution as introduced in \citet{li2019jan}. The MGL copula can capture nonelliptical, exchangeable, and asymmetric dependencies among marginal coordinates and provides a simple formulation for regression applications. We discuss the probabilistic characteristics of MGL copula and obtain the corresponding extreme-value copula, named the MGL-EV copula. While the survival MGL copula can be also regarded as a special case of the MGB2 copula from \citet{yang2011generalized}, we show that the proposed model is effective in regression modelling of dependence structures. Next to a simulation study, we propose two applications illustrating the usefulness of the proposed model. This method is also implemented in a user-friendly R package: \texttt{rMGLReg}.

stat.ME

Trimmed extreme value estimators for censored heavy-tailed data

We consider estimation of the extreme value index and extreme quantiles for heavy-tailed data that are right-censored. We study a general procedure of removing low importance observations in tail estimators. This trimming procedure is applied to the state-of-the-art estimators for randomly right-censored tail estimators. Through an averaging procedure over the amount of trimming we derive new kernel type estimators. Extensive simulation suggests that one of the new considered kernels leads to a highly competitive estimator against virtually any other available alternative in this framework. Moreover, we propose an adaptive selection method for the amount of top data used in estimation based on the trimming procedure minimizing the asymptotic mean squared error. We also provide an illustration of this approach to simulated as well as to real-world MTPL insurance data.

math.ST

Tempered Pareto-type modelling using Weibull distributions

In various applications of heavy-tail modelling, the assumed Pareto behavior is tempered ultimately in the range of the largest data. In insurance applications, claim payments are influenced by claim management and claims may for instance be subject to a higher level of inspection at highest damage levels leading to weaker tails than apparent from modal claims. Generalizing earlier results of Meerschaert et al. (2012) and Raschke (2019), in this paper we consider tempering of a Pareto-type distribution with a general Weibull distribution in a peaks-over-threshold approach. This requires to modulate the tempering parameters as a function of the chosen threshold. Modelling such a tempering effect is important in order to avoid overestimation of risk measures such as the Value-at-Risk (VaR) at high quantiles. We use a pseudo maximum likelihood approach to estimate the model parameters, and consider the estimation of extreme quantiles. We derive basic asymptotic results for the estimators, give illustrations with simulation experiments and apply the developed techniques to fire and liability insurance data, providing insight into the relevance of the tempering component in heavy-tail modelling.

math.ST

Generalizing the log-Moyal distribution and regression models for heavy tailed loss data

Catastrophic loss data are known to be heavy-tailed. Practitioners then need models that are able to capture both tail and modal parts of claim data. To this purpose, a new parametric family of loss distributions is proposed as a gamma mixture of the generalized log-Moyal distribution from Bhati and Ravi (2018), termed the generalized log-Moyal gamma distribution (GLMGA). We discuss the probabilistic characteristics of the GLMGA, and statistical estimation of the parameters through maximum likelihood. While the GLMGA distribution is a special case of the GB2 distribution, we show that this simpler model is effective in regression modelling of large and modal loss data. A fire claim data set reported in Cummins et al. (1990) and a Chinese earthquake loss data set are used to illustrate the applicability of the proposed model.

stat.AP

Center-outward quantiles and the measurement of multivariate risk

All multivariate extensions of the univariate theory of risk measurement run into the same fundamental problem of the absence, in dimension d > 1, of a canonical ordering of Rd. Based on measure transportation ideas, several attempts have been made recently in the statistical literature to overcome that conceptual difficulty. In Hallin (2017), the concepts of center-outward distribution and quantile functions are developed as generalisations of the classical univariate concepts of distribution and quantile functions, along with their empirical versions. We propose a class of smooth approximations as an alternative to the interpolation developed in del Barrio et al. (2018). This approximation allows for the computation of some new empirical risk measures, based either on the convex potential associated with the proposed transports, or on the volumes of the resulting empirical quantile regions. We also discuss the role of such transports in the evaluation of the risk associated with multivariate regularly varying distributions. Some simulations and applications to case studies illustrate the value of the approach.

stat.ME

Outlier detection and a tail-adjusted boxplot based on extreme value theory

Whether an extreme observation is an outlier or not, depends strongly on the corresponding tail behaviour of the underlying distribution. We develop an automatic, data-driven method to identify extreme tail behaviour that deviates from the intermediate and central characteristics. This allows for detecting extreme outliers or sets of extreme data that show less spread than the bulk of the data. To this end we extend a testing method proposed in Bhattacharya et al 2019 for the specific case of heavy tailed models, to all max-domains of attraction. Consequently we propose a tail-adjusted boxplot which yields a more accurate representation of possible outliers. Several examples and simulation results illustrate the finite sample behaviour of this approach.

stat.ME

Combined Tail Estimation Using Censored Data and Expert Information

We study tail estimation in Pareto-like settings for datasets with a high percentage of randomly right-censored data, and where some expert information on the tail index is available for the censored observations. This setting arises for instance naturally for liability insurance claims, where actuarial experts build reserves based on the specificity of each open claim, which can be used to improve the estimation based on the already available data points from closed claims. Through an entropy-perturbed likelihood we derive an explicit estimator and establish a close analogy with Bayesian methods. Embedded in an extreme value approach, asymptotic normality of the estimator is shown, and when the expert is clair-voyant, a simple combination formula can be deduced, bridging the classical statistical approach with the expert information. Following the aforementioned combination formula, a combination of quantile estimators can be naturally defined. In a simulation study, the estimator is shown to often outperform the Hill estimator for censored observations and recent Bayesian solutions, some of which require more information than usually available. Finally we perform a case study on a motor third-party liability insurance claim dataset, where Hill-type and quantile plots incorporate ultimate values into the estimation procedure in an intuitive manner.

stat.AP

Threshold selection and trimming in extremes

We consider removing lower order statistics from the classical Hill estimator in extreme value statistics, and compensating for it by rescaling the remaining terms. Trajectories of these trimmed statistics as a function of the extent of trimming turn out to be quite flat near the optimal threshold value. For the regularly varying case, the classical threshold selection problem in tail estimation is then revisited, both visually via trimmed Hill plots and, for the Hall class, also mathematically via minimizing the expected empirical variance. This leads to a simple threshold selection procedure for the classical Hill estimator which circumvents the estimation of some of the tail characteristics, a problem which is usually the bottleneck in threshold selection. As a by-product, we derive an alternative estimator of the tail index, which assigns more weight to large observations, and works particularly well for relatively lighter tails. A simple ratio statistic routine is suggested to evaluate the goodness of the implied selection of the threshold. We illustrate the favourable performance and the potential of the proposed method with simulation studies and real insurance data.

stat.ME

Bias Reduced Peaks over Threshold Tail Estimation

In recent years several attempts have been made to extend tail modelling towards the modal part of the data. Frigessi et al. (2002) introduced dynamic mixtures of two components with a weight function {\pi} = {\pi}(x) smoothly connecting the bulk and the tail of the distribution. Recently, Naveau et al. (2016) reviewed this topic, and, continuing on the work by Papastathopoulos and Tawn (2013), proposed a statistical model which is in compliance with extreme value theory and allows for a smooth transition between the modal and tail part. Incorporating second order rates of convergence for distributions of peaks over thresholds (POT), Beirlant et al. (2002, 2009) constructed models that can be viewed as special cases from both approaches discussed above. When fitting such second order models it turns out that the bias of the resulting extreme value estimators is significantly reduced compared to the classical tail fits using only the first order tail component based on the Pareto or generalized Pareto fits to peaks over threshold distributions. In this paper we provide novel bias reduced tail fitting techniques, improving upon the classical generalized Pareto (GP) approximation for POTs using the flexible semiparametric GP modelling introduced in Tencaliec et al. (2018). We also revisit and extend the secondorder refined POT approach started in Beirlant et al. (2009) to all max-domains of attraction using flexible semiparametric modelling of the second order component. In this way we relax the classical second order regular variation assumptions.

stat.ME

Estimation of the extreme value index in a censorship framework: asymptotic and finite sample behaviour

We revisit the estimation of the extreme value index for randomly censored data from a heavy tailed distribution. We introduce a new class of estimators which encompasses earlier proposals given in Worms and Worms (2014) and Beirlant et al. (2018), which were shown to have good bias properties compared with the pseudo maximum likelihood estimator proposed in Beirlant et al. (2007) and Einmahl et al. (2008). However the asymptotic normality of the type of estimators first proposed in Worms and Worms (2014) was still lacking, in the random threshold case. We derive an asymptotic representation and the asymptotic normality of the larger class of estimators and consider their finite sample behaviour. Special attention is paid to the case of heavy censoring, i.e. where the amount of censoring in the tail is at least 50\%. We obtain the asymptotic normality with a classical $\sqrt{k}$ rate where $k$ denotes the number of top data used in the estimation, depending on the degree of censoring.

math.ST

Estimating the maximum possible earthquake magnitude using extreme value methodology: the Groningen case

The area-characteristic, maximum possible earthquake magnitude $T_M$ is required by the earthquake engineering community, disaster management agencies and the insurance industry. The Gutenberg-Richter law predicts that earthquake magnitudes $M$ follow a truncated exponential distribution. In the geophysical literature several estimation procedures were proposed, see for instance Kijko and Singh (Acta Geophys., 2011) and the references therein. Estimation of $T_M$ is of course an extreme value problem to which the classical methods for endpoint estimation could be applied. We argue that recent methods on truncated tails at high levels (Beirlant et al., Extremes, 2016; Electron. J. Stat., 2017) constitute a more appropriate setting for this estimation problem. We present upper confidence bounds to quantify uncertainty of the point estimates. We also compare methods from the extreme value and geophysical literature through simulations. Finally, the different methods are applied to the magnitude data for the earthquakes induced by gas extraction in the Groningen province of the Netherlands.

stat.AP

Penalized bias reduction in extreme value estimation for censored Pareto-type data, and long-tailed insurance applications

The subject of tail estimation for randomly censored data from a heavy tailed distribution receives growing attention, motivated by applications for instance in actuarial statistics. The bias of the available estimators of the extreme value index can be substantial and depends strongly on the amount of censoring. We review the available estimators, propose a new bias reduced estimator, and show how shrinkage estimation can help to keep the MSE under control. A bootstrap algorithm is proposed to construct confidence intervals. We compare these new proposals with the existing estimators through simulation. We conclude this paper with a detailed study of a long-tailed car insurance portfolio, which typically exhibit heavy censoring.

stat.ME

Modelling Censored Losses Using Splicing: a Global Fit Strategy With Mixed Erlang and Extreme Value Distributions

In risk analysis, a global fit that appropriately captures the body and the tail of the distribution of losses is essential. Modelling the whole range of the losses using a standard distribution is usually very hard and often impossible due to the specific characteristics of the body and the tail of the loss distribution. A possible solution is to combine two distributions in a splicing model: a light-tailed distribution for the body which covers light and moderate losses, and a heavy-tailed distribution for the tail to capture large losses. We propose a splicing model with a mixed Erlang (ME) distribution for the body and a Pareto distribution for the tail. This combines the flexibility of the ME distribution with the ability of the Pareto distribution to model extreme values. We extend our splicing approach for censored and/or truncated data. Relevant examples of such data can be found in financial risk analysis. We illustrate the flexibility of this splicing model using practical examples from risk measurement.

stat.ME

Reducing MSE in estimation of heavy tails: a Bayesian approach

Bias reduction in tail estimation has received considerable interest in extreme value analysis. Estimation methods that minimize the bias while keeping the mean squared error (MSE) under control, are especially useful when applying classical methods such as the Hill (1975) estimator. In Caeiro et al. (2005) minimum variance reduced bias estimators of the Pareto tail index were first proposed where the bias is reduced without increasing the variance with respect to the Hill estimator. This method is based on adequate external estimation of a pair of second-order parameters. Here we revisit this problem from a Bayesian point of view starting from the extended Pareto distribution (EPD) approximation to excesses over a high threshold, as developed in Beirlant et al. (2009) using maximum likelihood (ML) estimation. Using asymptotic considerations, we derive an appropriate choice of priors leading to a Bayes estimator for which the MSE curve is a weighted average of the Hill and EPD-ML MSE curves for a large range of thresholds, under the same conditions as in Beirlant et al.(2009). A similar result is obtained for tail probability estimation. Simulations show surprisingly good MSE performance with respect to the existing estimators.

math.ST

Fitting tails affected by truncation

In several applications, ultimately at the largest data, truncation effects can be observed when analysing tail characteristics of statistical distributions. In some cases truncation effects are forecasted through physical models such as the Gutenberg-Richter relation in geophysics, while at other instances the nature of the measurement process itself may cause under recovery of large values, for instance due to flooding in river discharge readings. Recently Beirlant et al. (2016) discussed tail fitting for truncated Pareto-type distributions. Using examples from earthquake analysis, hydrology and diamond valuation we demonstrate the need for a unified treatment of extreme value analysis for truncated heavy and light tails. We generalise the classical Peaks over Threshold approach for the different max-domains of attraction with shape parameter $\xi>-1/2$ to allow for truncation effects. We use a pseudo-maximum likelihood approach to estimate the model parameters and consider extreme quantile estimation and reconstruction of quantile levels before truncation whenever appropriate. We report on some simulation experiments and provide some basic asymptotic results.

stat.ME

Tail fitting for truncated and non-truncated Pareto-type distributions

Recently some papers, such as Aban, Meerschaert and Panorska (2006), Nuyts (2010) and Clark (2013), have drawn attention to possible truncation in Pareto tail modelling. Sometimes natural upper bounds exist that truncate the probability tail, such as the Maximum Possible Loss in insurance treaties. At other instances ultimately at the largest data, deviations from a Pareto tail behaviour become apparent. This matter is especially important when extrapolation outside the sample is required. Given that in practice one does not always know whether the distribution is truncated or not, we consider estimators for extreme quantiles both under truncated and non-truncated Pareto-type distributions. Hereby we make use of the estimator of the tail index for the truncated Pareto distribution first proposed in Aban {\it et al.} (2006). We also propose a truncated Pareto QQ-plot and a formal test for truncation in order to help deciding between a truncated and a non-truncated case. In this way we enlarge the possibilities of extreme value modelling using Pareto tails, offering an alternative scenario by adding a truncation point $T$ that is large with respect to the available data. In the mathematical modelling we hence let $T \to \infty$ at different speeds compared to the limiting fraction ($k/n \to 0$) of data used in the extreme value estimation. This work is motivated using practical examples from different fields of applications, simulation results, and some asymptotic results.

math.ST