SearcharxivSearch

arXiv subjects

Mathias Lindholm

Publications and source records attributed to Mathias Lindholm.

14 recordsLinked to original sources

Forecasting sub-population mortality using credibility theory

The focus of the present paper is to forecast mortality rates for small sub-populations that are parts of a larger super-population. In this setting the assumption is that it is possible to produce reliable forecasts for the super-population, but the sub-populations may be too small or lack sufficient history to produce reliable forecasts if modelled separately. This setup is aligned with the ideas that underpin credibility theory, and in the present paper the classical credibility theory approach is extended to be able to handle the situation where future mortality rates are driven by a latent stochastic process, as is the case for, e.g., Lee-Carter type models. This results in sub-population credibility predictors that are weighted averages of expected future super-population mortality rates and expected future sub-population specific mortality rates. Due to the predictor's simple structure it is possible to derive an explicit expression for the expected quadratic forecast error. Moreover, the proposed credibility modelling approach does not depend on the specific form of the super-population model, making it broadly applicable regardless of the chosen forecasting model for the super-population. The performance of the suggested sub-population credibility predictor is illustrated on simulated population data. These illustrations highlight how the credibility predictor serves as a compromise between only using a super-population model, and only using a potentially unreliable sub-population specific model.

stat.AP

Regularisation of CART trees by summation of $p$-values

The standard procedure to decide on the complexity of a CART regression tree is to use cross-validation with the aim of obtaining a predictor that generalises well to unseen data. The randomness in the selection of folds implies that the selected CART regression tree is not a deterministic function of the data. Moreover, the cross-validation procedure may become time consuming and result in inefficient use of training data. We propose a simple deterministic in-sample method that can be used for stopping the growing of a CART regression tree based on node-wise statistical tests. This testing procedure is derived using a connection to change point detection, where the null hypothesis corresponds to no signal. The suggested $p$-value based procedure allows us to consider covariate vectors of arbitrary dimension and allows us to bound the $p$-value of an entire tree from above. Further, we show that the test detects a not too weak signal with a high probability, given a not too small sample size. We illustrate our methodology and the asymptotic results on both simulated and real world data. Additionally, we illustrate how the $p$-value based method can be used to construct a deterministic piece-wise constant auto-calibrated predictor based on a given black-box predictor.

stat.ME

A tree-based varying coefficient model

The paper introduces a tree-based varying coefficient model (VCM) where the varying coefficients are modelled using the cyclic gradient boosting machine (CGBM) from Delong et al. (2023). Modelling the coefficient functions using a CGBM allows for dimension-wise early stopping and feature importance scores. The dimension-wise early stopping not only reduces the risk of dimension-specific overfitting, but also reveals differences in model complexity across dimensions. The use of feature importance scores allows for simple feature selection and easy model interpretation. The model is evaluated on the same simulated and real data examples as those used in Richman and W\"uthrich (2023), and the results show that it produces results in terms of out of sample loss that are comparable to those of their neural network-based VCM called LocalGLMnet.

stat.ML

Mortality Forecasting using Variational Inference

This paper considers the problem of forecasting mortality rates. A large number of models have already been proposed for this task, but they generally have the disadvantage of either estimating the model in a two-step process, possibly losing efficiency, or relying on methods that are cumbersome for the practitioner to use. We instead propose using variational inference and the probabilistic programming library Pyro for estimating the model. This allows for flexibility in modelling assumptions while still being able to estimate the full model in one step. The models are fitted on Swedish mortality data and we find that the in-sample fit is good and that the forecasting performance is better than other popular models. Code is available at https://github.com/LPAndersson/VImortality.

stat.AP

A Discussion of Discrimination and Fairness in Insurance Pricing

Indirect discrimination is an issue of major concern in algorithmic models. This is particularly the case in insurance pricing where protected policyholder characteristics are not allowed to be used for insurance pricing. Simply disregarding protected policyholder information is not an appropriate solution because this still allows for the possibility of inferring the protected characteristics from the non-protected ones. This leads to so-called proxy or indirect discrimination. Though proxy discrimination is qualitatively different from the group fairness concepts in machine learning, these group fairness concepts are proposed to 'smooth out' the impact of protected characteristics in the calculation of insurance prices. The purpose of this note is to share some thoughts about group fairness concepts in the light of insurance pricing and to discuss their implications. We present a statistical model that is free of proxy discrimination, thus, unproblematic from an insurance pricing point of view. However, we find that the canonical price in this statistical model does not satisfy any of the three most popular group fairness axioms. This seems puzzling and we welcome feedback on our example and on the usefulness of these group fairness axioms for non-discriminatory insurance pricing.

cs.LG

A multi-task network approach for calculating discrimination-free insurance prices

In applications of predictive modeling, such as insurance pricing, indirect or proxy discrimination is an issue of major concern. Namely, there exists the possibility that protected policyholder characteristics are implicitly inferred from non-protected ones by predictive models, and are thus having an undesirable (or illegal) impact on prices. A technical solution to this problem relies on building a best-estimate model using all policyholder characteristics (including protected ones) and then averaging out the protected characteristics for calculating individual prices. However, such approaches require full knowledge of policyholders' protected characteristics, which may in itself be problematic. Here, we address this issue by using a multi-task neural network architecture for claim predictions, which can be trained using only partial information on protected characteristics, and it produces prices that are free from proxy discrimination. We demonstrate the use of the proposed model and we find that its predictive accuracy is comparable to a conventional feedforward neural network (on full information). However, this multi-task network has clearly superior performance in the case of partially missing policyholder information.

cs.LG

Bayesian Quantile-Based Portfolio Selection

We study the optimal portfolio allocation problem from a Bayesian perspective using value at risk (VaR) and conditional value at risk (CVaR) as risk measures. By applying the posterior predictive distribution for the future portfolio return, we derive relevant quantiles needed in the computations of VaR and CVaR, and express the optimal portfolio weights in terms of observed data only. This is in contrast to the conventional method where the optimal solution is based on unobserved quantities which are estimated, leading to suboptimality. We also obtain the expressions for the weights of the global minimum VaR and CVaR portfolios, and specify conditions for their existence. It is shown that these portfolios may not exist if the confidence level used for the VaR or CVaR computation are too low. Moreover, analytical expressions for the mean-VaR and mean-CVaR efficient frontiers are presented and the extension of theoretical results to general coherent risk measures is provided. One of the main advantages of the suggested Bayesian approach is that the theoretical results are derived in the finite-sample case and thus they are exact and can be applied to large-dimensional portfolios. By using simulation and real market data, we compare the new Bayesian approach to the conventional method by studying the performance and existence of the global minimum VaR portfolio and by analysing the estimated efficient frontiers. It is concluded that the Bayesian approach outperforms the conventional one, in particular at predicting the out-of-sample VaR.

q-fin.PM

How to ask sensitive multiple choice questions

Motivated by recent failures of polling to estimate populist party support, we propose and analyse two methods for asking sensitive multiple choice questions where the respondent retains some privacy and therefore might answer more truthfully. The first method consists of asking for the true choice along with a choice picked at random. The other method presents a list of choices and asks whether the preferred one is on the list or not. Different respondents are shown different lists. The methods are easy to explain, which makes it likely that the respondent understands how her privacy is protected and may thus entice her to participate in the survey and answer truthfully. The methods are also easy to implement and scale up.

stat.ME

Insurance valuation: a computable multi-period cost-of-capital approach

We present an approach to market-consistent multi-period valuation of insurance liability cash flows based on a two-stage valuation procedure. First, a portfolio of traded financial instrument aimed at replicating the liability cash flow is fixed. Then the residual cash flow is managed by repeated one-period replication using only cash funds. The latter part takes capital requirements and costs into account, as well as limited liability and risk averseness of capital providers. The cost-of-capital margin is the value of the residual cash flow. We set up a general framework for the cost-of-capital margin and relate it to dynamic risk measurement. Moreover, we present explicit formulas and properties of the cost-of-capital margin under further assumptions on the model for the liability cash flow and on the conditional risk measures and utility functions. Finally, we highlight computational aspects of the cost-of-capital margin, and related quantities, in terms of an example from life insurance.

q-fin.RM

Issues with the Smith-Wilson method

The objective of the present paper is to analyse various features of the Smith-Wilson method used for discounting under the EU regulation Solvency II, with special attention to hedging. In particular, we show that all key rate duration hedges of liabilities beyond the Last Liquid Point will be peculiar. Moreover, we show that there is a connection between the occurrence of negative discount factors and singularities in the convergence criterion used to calibrate the model. The main tool used for analysing hedges is a novel stochastic representation of the Smith-Wilson method. Further, we provide necessary conditions needed in order to construct similar, but hedgeable, discount curves.

q-fin.PR

Growing networks with preferential addition and deletion of edges

A preferential attachment model for a growing network incorporating deletion of edges is studied and the expected asymptotic degree distribution is analyzed. At each time step $t=1,2,\ldots$, with probability $π_1>0$ a new vertex with one edge attached to it is added to the network and the edge is connected to an existing vertex chosen proportionally to its degree, with probability $π_2$ a vertex is chosen proportionally to its degree and an edge is added between this vertex and a randomly chosen other vertex, and with probability $π_3=1-π_1-π_2<1/2$ a vertex is chosen proportionally to its degree and a random edge of this vertex is deleted. The model is intended to capture a situation where high-degree vertices are more dynamic than low-degree vertices in the sense that their connections tend to be changing. A recursion formula is derived for the expected asymptotic fraction $p_k$ of vertices with degree $k$, and solving this recursion reveals that, for $π_3<1/3$, we have $p_k\sim k^{-(3-7π_3)/(1-3π_3)}$, while, for $π_3>1/3$, the fraction $p_k$ decays exponentially at rate $(π_1+π_2)/2π_3$. There is hence a non-trivial upper bound for how much deletion the network can incorporate without loosing the power-law behavior of the degree distribution. The analytical results are supported by simulations.

physics.soc-ph

A dynamic network in a dynamic population: asymptotic properties

We derive asymptotic properties for a stochastic dynamic network model in a stochastic dynamic population. In the model, nodes give birth to new nodes until they die, each node being equipped with a social index given at birth. During the life of a node it creates edges to other nodes, nodes with high social index at higher rate, and edges disappear randomly in time. For this model we derive criterion for when a giant connected component exists after the process has evolved for a long period of time, assuming the node population grows to infinity. We also obtain an explicit expression for the degree correlation $ρ$ (of neighbouring nodes) which shows that $ρ$ is always positive irrespective of parameter values in one of the two treated submodels, and may be either positive or negative in the other model, depending on the parameters.

math.PR

A note on the component structure in random intersection graphs with tunable clustering

We study the component structure in random intersection graphs with tunable clustering, and show that the average degree works as a threshold for a phase transition for the size of the largest component. That is, if the expected degree is less than one, the size of the largest component is a.a.s. of logarithmic order, but if the average degree is greater than one, a.a.s. a single large component of linear order emerges, and the size of the second largest component is at most of logarithmic order.

math.PR

Epidemics on random graphs with tunable clustering

In this paper, a branching process approximation for the spread of a Reed-Frost epidemic on a network with tunable clustering is derived. The approximation gives rise to expressions for the epidemic threshold and the probability of a large outbreak in the epidemic. It is investigated how these quantities varies with the clustering in the graph and it turns out for instance that, as the clustering increases, the epidemic threshold decreases. The network is modelled by a random intersection graph, in which individuals are independently members of a number of groups and two individuals are linked to each other if and only if they share at least one group.

math.PR