SearcharxivSearch

arXiv subjects

Gareth W. Peters

Publications and source records attributed to Gareth W. Peters.

At least 19 recordsLinked to original sources

Market-based insurance ratemaking: application to pet insurance

This paper introduces a method for pricing insurance policies using market data. The approach is designed for scenarios in which the insurance company seeks to enter a new market, in our case: pet insurance, lacking historical data. The methodology involves an iterative two-step process. First, a suitable parameter is proposed to characterize the underlying risk. Second, the resulting pure premium is linked to the observed commercial premium using an isotonic regression model. To validate the method, comprehensive testing is conducted on synthetic data, followed by its application to a dataset of actual pet insurance rates. To facilitate practical implementation, we have developed an R package called IsoPriceR. By addressing the challenge of pricing insurance policies in the absence of historical data, this method helps enhance pricing strategies in emerging markets.

stat.AP

PDSim: A Shiny App for Simulating and Estimating Polynomial Diffusion Models in Commodity Futures

PDSim is an R package that enables users to simulate commodity futures prices using the polynomial diffusion model introduced in Filipovic & Larsson (2016) through both a Shiny web application and R scripts. For user-supplied data, a standalone R routine has been developed to provide joint estimation of state variables and model parameters via the Extended Kalman Filter (EKF) or Unscented Kalman Filter (UKF). With its user-friendly interface, PDSim makes the features of simulations and estimations accessible. To date, it is the only package specifically designed for the simulation and estimation of the polynomial diffusion model. The Schwartz-Smith two-factor model (Schwartz & Smith, 2000) is also available within this package for both simulation and calibration. The package is validated through several tests, including replication of the results in Schwartz & Smith (2000), unit testing of the coverage rate, and verification of the outputs of the main functions.

q-fin.ST

Signature Isolation Forest

Functional Isolation Forest (FIF) is a recent state-of-the-art Anomaly Detection (AD) algorithm designed for functional data. It relies on a tree partition procedure where an abnormality score is computed by projecting each curve observation on a drawn dictionary through a linear inner product. Such linear inner product and the dictionary are a priori choices that highly influence the algorithm's performances and might lead to unreliable results, particularly with complex datasets. This work addresses these challenges by introducing \textit{Signature Isolation Forest}, a novel AD algorithm class leveraging the rough path theory's signature transform. Our objective is to remove the constraints imposed by FIF through the proposition of two algorithms which specifically target the linearity of the FIF inner product and the choice of the dictionary. We provide several numerical experiments, including a real-world applications benchmark showing the relevance of our methods.

stat.ML

Multi-Factor Function-on-Function Regression of Bond Yields on WTI Commodity Futures Term Structure Dynamics

In the analysis of commodity futures, it is commonly assumed that futures prices are driven by two latent factors: short-term fluctuations and long-term equilibrium price levels. In this study, we extend this framework by introducing a novel state-space functional regression model that incorporates yield curve dynamics. Our model offers a distinct advantage in capturing the interdependencies between commodity futures and the yield curve. Through a comprehensive empirical analysis of WTI crude oil futures, using US Treasury yields as a functional predictor, we demonstrate the superior accuracy of the functional regression model compared to the Schwartz-Smith two-factor model, particularly in estimating the short-end of the futures curve. Additionally, we conduct a stress testing analysis to examine the impact of both temporary and permanent shocks to US Treasury yields on futures price estimation.

q-fin.ST

Cyber Risk Taxonomies: Statistical Analysis of Cybersecurity Risk Classifications

Cyber risk classifications are widely used in the modeling of cyber event distributions, yet their effectiveness in out of sample forecasting performance remains underexplored. In this paper, we analyse the most commonly used classifications and argue in favour of switching the attention from goodness-of-fit and in-sample predictive performance, to focusing on the out-of sample forecasting performance. We use a rolling window analysis, to compare cyber risk distribution forecasts via threshold weighted scoring functions. Our results indicate that business motivated cyber risk classifications appear to be too restrictive and not flexible enough to capture the heterogeneity of cyber risk events. We investigate how dynamic and impact-based cyber risk classifiers seem to be better suited in forecasting future cyber risk losses than the other considered classifications. These findings suggest that cyber risk types provide limited forecasting ability concerning cyber event severity distribution, and cyber insurance ratemakers should utilize cyber risk types only when modeling the cyber event frequency distribution. Our study offers valuable insights for decision-makers and policymakers alike, contributing to the advancement of scientific knowledge in the field of cyber risk management.

cs.CR

Multi-Factor Polynomial Diffusion Models and Inter-Temporal Futures Dynamics

In stochastic multi-factor commodity models, it is often the case that futures prices are explained by two latent state variables which represent the short and long term stochastic factors. In this work, we develop the family of stochastic models using polynomial diffusion to obtain the unobservable spot price to be used for modelling futures curve dynamics. The polynomial family of diffusion models allows one to incorporate a variety of non-linear, higher-order effects, into a multi-factor stochastic model, which is a generalisation of Schwartz and Smith (2000) two-factor model. Two filtering methods are used for the parameter and the latent factor estimation to address the non-linearity. We provide a comparative analysis of the performance of the estimation procedures. We discuss the parameter identification problem present in the polynomial diffusion case, regardless, the futures prices can still be estimated accurately. Moreover, we study the effects of different methods of calculating matrix exponential in the polynomial diffusion model. As the polynomial order increases, accurately and efficiently approximating the high-dimensional matrix exponential becomes essential in the polynomial diffusion model.

q-fin.ST

State-Space Dynamic Functional Regression for Multicurve Fixed Income Spread Analysis and Stress Testing

The Nelson-Siegel model is widely used in fixed income markets to produce yield curve dynamics. The multiple time-dependent parameter model conveniently addresses the level, slope, and curvature dynamics of the yield curves. In this study, we present a novel state-space functional regression model that incorporates a dynamic Nelson-Siegel model and functional regression formulations applied to multi-economy setting. This framework offers distinct advantages in explaining the relative spreads in yields between a reference economy and a response economy. To address the inherent challenges of model calibration, a kernel principal component analysis is employed to transform the representation of functional regression into a finite-dimensional, tractable estimation problem. A comprehensive empirical analysis is conducted to assess the efficacy of the functional regression approach, including an in-sample performance comparison with the dynamic Nelson-Siegel model. We conducted the stress testing analysis of yield curves term-structure within a dual economy framework. The bond ladder portfolio was examined through a case study focused on spread modelling using historical data for US Treasury and UK bonds.

q-fin.ST

A Bonus-Malus Framework for Cyber Risk Insurance and Optimal Cybersecurity Provisioning

The cyber risk insurance market is at a nascent stage of its development, even as the magnitude of cyber losses is significant and the rate of cyber loss events is increasing. Existing cyber risk insurance products as well as academic studies have been focusing on classifying cyber loss events and developing models of these events, but little attention has been paid to proposing insurance risk transfer strategies that incentivise mitigation of cyber loss through adjusting the premium of the risk transfer product. To address this important gap, we develop a Bonus-Malus model for cyber risk insurance. Specifically, we propose a mathematical model of cyber risk insurance and cybersecurity provisioning supported with an efficient numerical algorithm based on dynamic programming. Through a numerical experiment, we demonstrate how a properly designed cyber risk insurance contract with a Bonus-Malus system can resolve the issue of moral hazard and benefit the insurer.

math.OC

Cyber Loss Model Risk Translates to Premium Mispricing and Risk Sensitivity

We focus on model risk and risk sensitivity when addressing the insurability of cyber risk. The standard statistical approaches to assessment of insurability and potential mispricing are enhanced in several aspects involving consideration of model risk. Model risk can arise from model uncertainty, and parameters uncertainty. We demonstrate how to quantify the effect of model risk in this analysis by incorporating various robust estimators for key model parameter estimates that apply in both marginal and joint cyber risk loss process modelling. We contrast these robust techniques with standard methods previously used in studying insurabilty of cyber risk. This allows us to accurately assess the critical impact that robust estimation can have on tail index estimation for heavy tailed loss models, as well as the effect of robust dependence analysis when quantifying joint loss models and insurance portfolio diversification. We argue that the choice of such methods is akin to a form of model risk and we study the risk sensitivity that arise from choices relating to the class of robust estimation adopted and the impact of the settings associated with such methods on key actuarial tasks such as premium calculation in cyber insurance. Through this analysis we are able to address the question that, to the best of our knowledge, no other study has investigated in the context of cyber risk: is model risk present in cyber risk data, and how does is it translate into premium mispricing? We believe our findings should complement existing studies seeking to explore insurability of cyber losses. In order to ensure our findings are based on realistic industry informed loss data, we have utilised one of the leading industry cyber loss datasets obtained from Advisen, which represents a comprehensive data set on cyber monetary losses, from which we form our analysis and conclusions.

q-fin.RM

Cyber Risk Frequency, Severity and Insurance Viability

In this study an exploration of insurance risk transfer is undertaken for the cyber insurance industry in the United States of America, based on the leading industry dataset of cyber events provided by Advisen. We seek to address two core unresolved questions. First, what factors are the most significant covariates that may explain the frequency and severity of cyber loss events and are they heterogeneous over cyber risk categories? Second, is cyber risk insurable in regards to the required premiums, risk pool sizes and how would this decision vary with the insured companies industry sector and size? We address these questions through a combination of regression models based on the class of Generalised Additive Models for Location Shape and Scale (GAMLSS) and a class of ordinal regressions. These models will then form the basis for our analysis of frequency and severity of cyber risk loss processes. We investigate the viability of insurance for cyber risk using a utility modelling framework with premium calculated by classical certainty equivalence analysis utilising the developed regression models. Our results provide several new key insights into the nature of insurability of cyber risk and rigorously address the two insurance questions posed in a real data driven case study analysis.

q-fin.RM

The Nature of Losses from Cyber-Related Events: Risk Categories and Business Sectors

In this study we examine the nature of losses from cyber related events across different risk categories and business sectors. Using a leading industry dataset of cyber events, we evaluate the relationship between the frequency and severity of individual cyber-related events and the number of affected records. We find that the frequency of reported cyber related events has substantially increased between 2008 and 2016. Furthermore, the frequency and severity of losses depend on the business sector and type of cyber threat: the most significant cyber loss event categories, by number of events, were related to data breaches and the unauthorized disclosure of data, while cyber extortion, phishing, spoofing and other social engineering practices showed substantial growth rates. Interestingly, we do not find a distinct pattern between the frequency of events, the loss severity, and the number of affected records as often alluded to in the literature. We also analyse the severity distribution of cyber related events across all risk categories and business sectors. This analysis reveals that cyber risks are heavy-tailed, i.e., cyber risk events have a higher probability to produce extreme losses than events whose severity follows an exponential distribution. Furthermore, we find that the frequency and severity of cyber related losses exhibits a very dynamic and time varying nature.

q-fin.RM

Stochastic measure distortions induced by quantile processes for risk quantification and valuation

We develop a novel stochastic valuation and premium calculation principle based on probability measure distortions that are induced by quantile processes in continuous time. Necessary and sufficient conditions are derived under which the quantile processes satisfy first- and second-order stochastic dominance. The introduced valuation principle relies on stochastic ordering so that the valuation risk-loading, and thus risk premiums, generated by the measure distortion is an ordered parametric family. The quantile processes are generated by a composite map consisting of a distribution and a quantile function. The distribution function accounts for model risk in relation to the empirical distribution of the risk process, while the quantile function models the response to the risk source as perceived by, e.g., a market agent. This gives rise to a system of subjective probability measures that indexes a stochastic valuation principle susceptible to probability measure distortions. We use the Tukey-$gh$ family of quantile processes driven by Brownian motion in an example that demonstrates stochastic ordering. We consider the conditional expectation under the distorted measure as a member of the time-consistent class of dynamic valuation principles, and extend it to the setting where the driving risk process is multivariate. This requires the introduction of a copula function in the composite map for the construction of quantile processes, which presents another new element in the risk quantification and modelling framework based on probability measure distortions induced by quantile processes.

q-fin.RM

Quantile Diffusions for Risk Analysis

We develop a novel approach for the construction of quantile processes governing the stochastic dynamics of quantiles in continuous time. Two classes of quantile diffusions are identified: the first, which we largely focus on, features a dynamic random quantile level and allows for direct interpretation of the resulting quantile process characteristics such as location, scale, skewness and kurtosis, in terms of the model parameters. The second type are function-valued quantile diffusions and are driven by stochastic parameter processes, which determine the entire quantile function at each point in time. By the proposed innovative and simple -- yet powerful -- construction method, quantile processes are obtained by transforming the marginals of a diffusion process under a composite map consisting of a distribution and a quantile function. Such maps, analogous to rank transmutation maps, produce the marginals of the resulting quantile process. We discuss the relationship and differences between our approach and existing methods and characterisations of quantile processes in discrete and continuous time. As an example of an application of quantile diffusions, we show how probability measure distortions, a form of dynamic tilting, can be induced. Though particularly useful in financial mathematics and actuarial science, examples of which are given in this work, measure distortions feature prominently across multiple research areas. For instance, dynamic distributional approximations (statistics), non-parametric and asymptotic analysis (mathematical statistics), dynamic risk measures (econometrics), behavioural economics, decision making (operations research), signal processing (information theory), and not least in general risk theory including applications thereof, for example in the context of climate change.

math.PR

Dynamic Quantile Function Models

Motivated by the need for effectively summarising, modelling, and forecasting the distributional characteristics of intra-daily returns, as well as the recent work on forecasting histogram-valued time-series in the area of symbolic data analysis, we develop a time-series model for forecasting quantile-function-valued (QF-valued) daily summaries for intra-daily returns. We call this model the dynamic quantile function (DQF) model. Instead of a histogram, we propose to use a $g$-and-$h$ quantile function to summarise the distribution of intra-daily returns. We work with a Bayesian formulation of the DQF model in order to make statistical inference while accounting for parameter uncertainty; an efficient MCMC algorithm is developed for sampling-based posterior inference. Using ten international market indices and approximately 2,000 days of out-of-sample data from each market, the performance of the DQF model compares favourably, in terms of forecasting VaR of intra-daily returns, against the interval-valued and histogram-valued time-series models. Additionally, we demonstrate that the QF-valued forecasts can be used to forecast VaR measures at the daily timescale via a simple quantile regression model on daily returns (QR-DQF). In certain markets, the resulting QR-DQF model is able to provide competitive VaR forecasts for daily returns.

stat.ME

Cost-aware Feature Selection for IoT Device Classification

Classification of IoT devices into different types is of paramount importance, from multiple perspectives, including security and privacy aspects. Recent works have explored machine learning techniques for fingerprinting (or classifying) IoT devices, with promising results. However, existing works have assumed that the features used for building the machine learning models are readily available or can be easily extracted from the network traffic; in other words, they do not consider the costs associated with feature extraction. In this work, we take a more realistic approach, and argue that feature extraction has a cost, and the costs are different for different features. We also take a step forward from the current practice of considering the misclassification loss as a binary value, and make a case for different losses based on the misclassification performance. Thereby, and more importantly, we introduce the notion of risk for IoT device classification. We define and formulate the problem of cost-aware IoT device classification. This being a combinatorial optimization problem, we develop a novel algorithm to solve it in a fast and effective way using the Cross-Entropy (CE) based stochastic optimization technique. Using traffic of real devices, we demonstrate the capability of the CE based algorithm in selecting features with minimal risk of misclassification while keeping the cost for feature extraction within a specified limit.

cs.NI

Multiple barrier-crossings of an Ornstein-Uhlenbeck diffusion in consecutive periods

We investigate the joint distribution and the multivariate survival functions for the maxima of an Ornstein-Uhlenbeck (OU) process in consecutive time-intervals. A PDE method, alongside an eigenfunction expansion, is adopted with which we first calculate the distribution and the survival functions for the maximum of a homogeneous OU-process in a single interval. By a deterministic time-change and a parameter translation, this result can be extended to an inhomogeneous OU-process. Next, we derive a general formula for the joint distribution and the survival functions for the maxima of a continuous Markov process in consecutive periods. With these results, one can obtain semi-analytical expressions for the joint distribution and the multivariate survival functions for the maxima of an OU-process, with piecewise constant parameter functions, in consecutive time periods. The joint distribution and the survival functions can be evaluated numerically by an iterated quadrature scheme, which can be implemented efficiently by matrix multiplications. Moreover, we show that the computation can be further simplified to the product of single quadratures by imposing a mild condition. Such results may be used for the modelling of heatwaves and related risk management challenges.

math.PR

Parsimonious Feature Extraction Methods: Extending Robust Probabilistic Projections with Generalized Skew-t

We propose a novel generalisation to the Student-t Probabilistic Principal Component methodology which: (1) accounts for an asymmetric distribution of the observation data; (2) is a framework for grouped and generalised multiple-degree-of-freedom structures, which provides a more flexible approach to modelling groups of marginal tail dependence in the observation data; and (3) separates the tail effect of the error terms and factors. The new feature extraction methods are derived in an incomplete data setting to efficiently handle the presence of missing values in the observation vector. We discuss various special cases of the algorithm being a result of simplified assumptions on the process generating the data. The applicability of the new framework is illustrated on a data set that consists of crypto currencies with the highest market capitalisation.

stat.ME

Bayesian Spatial Field Reconstruction with Unknown Distortions in Sensor Networks

Spatial regression of random fields based on potentially biased sensing information is proposed in this paper. One major concern in such applications is that since it is not known a-priori what the accuracy of the collected data from each sensor is, the performance can be negatively affected if the collected information is not fused appropriately. For example, the data collector may measure the phenomenon inappropriately, or alternatively, the sensors could be out of calibration, thus introducing random gain and bias to the measurement process. Such readings would be systematically distorted, leading to incorrect estimation of the spatial field. To combat this detrimental effect, we develop a robust version of the spatial field model based on a mixture of Gaussian process experts. We then develop two different approaches for Bayesian spatial field reconstruction: the first algorithm is the Spatial Best Linear Unbiased Estimator (S-BLUE), in which one considers the quadratic loss function and restricts the estimator to the linear family of transformations; the second algorithm is based on empirical Bayes, which utilises a two-stage estimation procedure to produce accurate predictive inference in the presence of "misbehaving" sensors. In addition, we develop the distributed version of these two approaches to drastically improve the computational efficiency in large-scale settings. We present extensive simulation results using both synthetic datasets and semi-synthetic datasets with real temperature measurements and simulated distortions to draw useful conclusions regarding the performance of each of the algorithms.

eess.SP