SearcharxivSearch

arXiv subjects

Christopher Blier-Wong

Publications and source records attributed to Christopher Blier-Wong.

18 recordsLinked to original sources

Towards foundation models for insurance risk modelling

Claim narratives, images and sensor data contain information about insured risks that is difficult to use through existing actuarial models. Foundation models learn patterns from large datasets before being adapted to particular tasks. By turning these high-dimensional sources into variables or numerical representations, they could help insurers use more of the information they already collect, potentially reducing the experience needed to develop each application. For example, a language model could identify a worsening injury in a new claim note, allowing a reserving model to recognise the change in expected cost before the payments reveal the deterioration. In this paper, we review language, vision, geospatial, time series, tabular and scientific models, explaining existing insurance applications and potential future uses. Scientific models extend this approach to future weather and climate conditions: their simulations can inform loss estimates once local hazards are linked to asset damage, repair costs and insurance coverage. We propose a process to connect these model outputs to actuarial calculations and to assess their predictive contribution, stability and compliance with rules on information use. Evaluating these applications is difficult when final claim costs become known only after long delays, large losses are rare or patterns learned elsewhere fail to transfer to the target portfolio. Richer data can reveal private information and support finer risk classification, which can change access to insurance. Reusing the same models across insurers also creates dependence on shared predictions and providers.

q-fin.RM

Sharp bounds for products of dependent random variables

We study the sharp bounds on $\mathbb{E}[X_1\cdots X_d]$ when the univariate marginal distributions are known, but the dependence structure between them is unspecified. Maximizing products over non-negative variables is straightforward via the comonotonic coupling, but the problem is more subtle when the marginals can take both positive and negative values. Specifically, two negative realizations can be matched to yield a positive product, whereas a single negative realization necessarily yields a negative product. We decompose the problem into a magnitude part and a sign part, and show that universal upper and lower bounds for the product expectation follow from comonotonic coupling of the absolute values and properly chosen sign vectors. Under a mild regularity assumption, we give necessary and sufficient conditions for these universal bounds to be attainable. For the upper bound, the marginal sign-bias vector must belong to the even-parity polytope, while for the lower, the corresponding condition involves the odd-parity polytope. We construct the extremal couplings via measurable selections on the parity polytope whenever these conditions hold. We study the case of identical marginals in more detail and provide examples of nonsymmetric extremal couplings that attain the universal bounds. We explicitly construct the extremal copulas in three dimensions, and use a recursive parity decomposition to obtain higher-dimensional extremal copulas from the trivariate ones.

math.ST

A multi-view contrastive learning framework for spatial embeddings in risk modelling

Incorporating spatial information, particularly when related to climate, weather, and demographic factors, is crucial for improving underwriting precision and enhancing risk management in insurance. However, spatial data are often unstructured, high-dimensional, and difficult to integrate into predictive models. Embedding methods are needed to convert spatial data into meaningful representations for modelling tasks. We propose a novel multi-view contrastive learning framework for generating spatial embeddings that combine information from multiple spatial data sources. To train the model, we construct a spatial dataset that merges satellite imagery and OpenStreetMap features across Europe. The framework aligns these spatial views with coordinate-based encodings, producing low-dimensional embeddings that capture both spatial structure and contextual similarity. Once trained, the model generates embeddings directly from latitude-longitude pairs, enabling any dataset with coordinates to be enriched with meaningful spatial features without requiring access to the original spatial inputs. In a case study on French real estate prices, we compare models trained on raw coordinates against those using our spatial embeddings as inputs. The embeddings consistently improve predictive accuracy across generalised linear, additive, and boosting models, while providing post-hoc explainable spatial effects and demonstrating generalisation of the fitted spatial effects to regions without training observations. A second case study on flood claim counts across Belgian postal codes confirms that the embeddings improve territorial risk classification in an insurance context.

q-fin.RM

Severity estimation in dependent collective risk models

The collective risk model represents the aggregate loss of an insurance portfolio as a random sum of individual claim severities. When claim counts and severities are dependent, the claims pooled across policies are no longer a sample from the marginal severity distribution. We show that their empirical distribution converges to the law of an arbitrary observed claim, a size-biased mixture of the conditional severity distributions, so any procedure that fits the severity margin directly to pooled claims is inconsistent in general. The same result identifies the distribution that the pooled claims do sample, and we build a composite likelihood estimation procedure on that distribution. We establish consistency and asymptotic normality, with Godambe information in which the policy, rather than the claim, is the sampling unit. In a Sarmanov collective risk model, the observed-claim density and the aggregate mean are in closed form. A simulation study measures the bias of naive pooled-severity fitting, its correction by the composite likelihood, and the coverage of the policy-level standard errors.

stat.ME

Semantic insurance pricing with large language models

Classical actuarial pricing models, such as the generalized linear model, are valued for transparency and ease of governance, but they use interactions among risk factors only when these are supplied through explicit feature engineering. We study whether embeddings from a pre-trained large language model, computed from a natural-language description of each policyholder, can replace hand-crafted features as inputs to a standard actuarial pricing model, taking Poisson claim-frequency regression as the main example. The language model is used only to construct deterministic embedding covariates; pricing is performed by a standard generalized linear model. Using French motor third-party liability data, the embedding-based model outperforms the generalized linear model, especially when data are scarce, whereas at larger sample sizes the comparison is model- and dimension-dependent. Insurance-specific fine-tuning further improves the embeddings, and a prompt-sensitivity diagnostic shows that the pipeline reacts to any appended out-of-template field, making controlled prompts a governance requirement.

stat.AP

Designing entry-monotone risk-sharing pools

While risk pooling lowers the total cost of risk, efficiency alone does not make a pool viable. Participants need terms that ensure their participation, that are immune to subgroups breaking away, and that allow new members to join. Under cash-additive risk measures, the minimum cost of a coalition's risk determines the value created by that coalition, and deterministic side payments redistribute that value among participants. Institutional risk sharing is thus a transferable-utility cooperative game. We prove that the game is totally balanced whenever the risk measures are convex (agents are risk averse), so every coalition has a nonempty core and stable allocations always exist. We then analyze entry monotonicity through Population-Monotonic Allocation Schemes (Sprumont, 1990), a strong requirement that is notoriously difficult to construct and has received limited attention in risk sharing. We find several structural conditions that ensure that either the Arrow--Debreu pricing surplus allocation rule or the proportional-cost surplus allocation rule satisfies this entry-monotonicity property, the latter being a novel cooperative notion we propose. These verifiable structural conditions naturally arise in pooled (re)insurance and credit portfolios, providing pool designers with a practical toolkit for building risk pools that remain stable and attractive as they expand.

econ.TH

Comonotonic improvement under feasibility constraints

Regulatory and contractual constraints on individual exposures are standard in insurance and reinsurance markets, but a poorly designed constraint can distort the economic incentives of risk-averse agents. In the unconstrained problem, the classical comonotonic improvement theorem guarantees Pareto-optimal allocations that are nondecreasing in the aggregate loss. A constraint that is not stable under risk reduction can destroy this property. We show by example that Value-at-Risk caps lead to optimal allocations that are non-comonotonic in the aggregate loss. We identify componentwise convex-order solidity as a sufficient condition on the feasible set that restores the comonotonic improvement under constraints. If replacing any agent's allocation by a less risky one preserves feasibility, then every feasible allocation admits a feasible comonotonic improvement for all convex-order-consistent preferences. This criterion covers many constraints typical in risk management, but excludes Value-at-Risk caps and idiosyncratic deductibles. We illustrate the implications of our main result in a mean-variance risk-sharing application.

econ.TH

A Laplace-based perspective on conditional mean risk sharing

The conditional mean risk-sharing (CMRS) rule is an important tool for distributing aggregate losses across individual risks, but its implementation in continuous multivariate models typically requires complicated multidimensional integrals. We develop a framework to compute CMRS allocations from the joint Laplace--Stieltjes transform of the risk vector. The LSTs of the allocation measures $ν_i(B)=\mathbb{E}[X_i\boldsymbol{1}_{\{S\in B\}}]$ are expressed as partial derivatives of the joint LST evaluated on the diagonal $t_1=\cdots=t_n$. When densities exist, this yields one-dimensional Laplace inversions for $f_S$ and $ξ_i$, and hence $h_i(s)=ξ_i(s)/f_S(s)$ on the absolutely continuous part, providing closed-form or semi-analytic solutions for a broad class of distributions. We also develop numerical inversion methods for cases where analytic inversion is unavailable. We introduce an exponential tilting procedure to stabilize numerical inversion in low-probability aggregate events. We provide several examples to illustrate the approach, including in some high-dimensional settings where existing approaches are infeasible.

math.ST

Stochastic representation of Sarmanov copulas

Sarmanov copulas offer a simple and tractable way to build multivariate distributions by perturbing the independence copula. They admit closed-form expressions for densities and many functionals of interest, making them attractive for practical applications. However, the complex conditions on the dependence parameters to ensure that Sarmanov copulas are valid limit their application in high dimensions. Verifying the $d$-increasing property typically requires satisfying a combinatorial set of inequalities that makes direct construction difficult. To circumvent this issue, we develop a stochastic representation for bivariate Sarmanov copulas. We prove that every admissible Sarmanov can be realized as a mixture of independent univariate distributions indexed by a latent Bernoulli pair. The stochastic representation replaces the problem of verifying copula validity with the problem of ensuring nonnegativity of a Bernoulli probability mass function. The representation also recovers classical copula families, including Farlie--Gumbel--Morgenstern, Huang--Kotz, and Bairamov--Kotz--Bekçi as special cases. We further derive sharp global bounds for Spearman's rho and Kendall's tau. We then introduce a Bernoulli-mixing construction in higher dimensions, leading to a new class of multivariate Sarmanov copulas with easily verifiable parameter constraints and scalable simulation algorithms. Finally, we show that powered versions of bivariate Sarmanov copulas admit a similar stochastic representation through block-maximal order statistics.

math.ST

Improved thresholds for e-values

The rejection threshold used for e-values and e-processes is by default set to $1/α$ for a guaranteed type-I error control at $α$, based on Markov's and Ville's inequalities. This threshold can be wasteful in practical applications. We discuss how this threshold can be improved under additional distributional assumptions on the e-values; some of these assumptions are naturally plausible and empirically observable, without knowing explicitly the form or model of the e-values. For small values of $α$, the threshold can roughly be improved (divided) by a factor of $2$ for decreasing or unimodal densities, and by a factor of $e$ for decreasing or unimodal-symmetric densities of log-transformed e-values. Moreover, we propose to use the supremum of comonotonic e-values, which is shown to preserve the type-I error guarantee. We also propose some preliminary methods to boost e-values in the e-BH procedure under some distributional assumptions while controlling the false discovery rate. Through a series of simulation studies, we demonstrate the effectiveness of our proposed methods in various testing scenarios, showing enhanced power.

math.ST

Efficient evaluation of risk allocations

Expectations of marginals conditional on the total risk of a portfolio are crucial in risk-sharing and allocation. However, computing these conditional expectations may be challenging, especially in critical cases where the marginal risks have compound distributions or when the risks are dependent. We introduce a generating function method to compute these conditional expectations. We provide efficient algorithms to compute the conditional expectations of marginals given the total risk for a portfolio of risks with lattice-type support. We show that the ordinary generating function of unconditional expected allocations is a function of the multivariate probability generating function of the portfolio. The generating function method allows us to develop recursive and transform-based techniques to compute the unconditional expected allocations. We illustrate our method to large-scale risk-sharing and risk allocation problems, including cases where the marginal risks have compound distributions, where the portfolio is composed of dependent risks, and where the risks have heavy tails, leading in some cases to computational gains of several orders of magnitude. Our approach is useful for risk-sharing in peer-to-peer insurance and risk allocation based on Euler's rule.

stat.AP

Collective risk models with FGM dependence

We study copula-based collective risk models when the dependence structure is defined by a Farlie-Gumbel-Morgenstern (FGM) copula. By leveraging a one-to-one correspondence between the class of FGM copulas and multivariate symmetric Bernoulli distributions, we find convenient representations for the moments and Laplace-Stieltjes transform for the aggregate random variable defined from collective risk models with FGM dependence. We examine different components of this collective risk model, aiming to better understand the impact of the assumed dependence between a claim's frequency and severity. Relying on stochastic ordering, we analyze the impact of dependence on the aggregate claim amount random variable. Even if the FGM copula may only induce moderate dependence, we illustrate through numerical examples that the cumulative effect of FGM dependence can lead to substantial variations in key risk measures on aggregate random variables defined from collective risk models.

stat.AP

A representation-learning approach for insurance pricing with images

Unstructured data are a promising new source of information that insurance companies may use to understand their risk portfolio better and improve the customer experience. However, these novel data sources are difficult to incorporate into existing ratemaking frameworks due to the size and format of the unstructured data. In this paper, we propose a framework to use street view imagery within a generalized linear model. To do so, we use representation learning to extract an embedding vector containing useful information from the image. This embedding is dense and low-dimensional, making it appropriate to use within existing ratemaking models. We find that there is useful information included in street view imagery to predict the frequency of claims for certain types of perils. This model can be used as-is in a ratemaking framework but also opens the door to future empirical research on attempting to extract the causal effect from images that lead to increased or decreased predicted claim frequencies. Throughout, we discuss the practical difficulties (technical and social) of using this type of data for insurance pricing.

stat.AP

A new method to construct high-dimensional copulas with Bernoulli and Coxian-2 distributions

We propose an approach to construct a new family of generalized Farlie-Gumbel-Morgenstern (GFGM) copulas that naturally scales to high dimensions. A GFGM copula can model moderate positive and negative dependence, cover different types of asymmetries, and admits exact expressions for many quantities of interest such as measures of association or risk measures in actuarial science or quantitative risk management. More importantly, this paper presents a new method to construct high-dimensional copulas based on mixtures of power functions, and may be adapted to more general contexts to construct broader families of copulas. We construct a family of copulas through a stochastic representation based on multivariate Bernoulli distributions and Coxian-2 distributions. This paper will cover the construction of a GFGM copula, study its measures of multivariate association and dependence properties. We explain how to sample random vectors from the new family of copulas in high dimensions. Then, we study the bivariate case in detail and find that our construction leads to an asymmetric modified Huang-Kotz FGM copula. Finally, we study the exchangeable case and provide some insights into the most negative dependence structure within this new class of high-dimensional copulas.

math.ST

Risk aggregation with FGM copulas

We offer a new perspective on risk aggregation with FGM copulas. Along the way, we discover new results and revisit existing ones, providing simpler formulas than one can find in the existing literature. This paper builds on two novel representations of FGM copulas based on symmetric multivariate Bernoulli distributions and order statistics. First, we detail families of multivariate distributions with closed-form solutions for the cumulative distribution function or moments of the aggregate random variables. We order aggregate random variables under the convex order and provide methods to compute the cumulative distribution function of aggregate rvs when the marginals are discrete. Finally, we discuss risk-sharing and capital allocation, providing numerical examples for each.

math.ST

Exchangeable FGM copulas

Copulas are a powerful tool to model dependence between the components of a random vector. One well-known class of copulas when working in two dimensions is the Farlie-GumbelMorgenstern (FGM) copula since their simple analytic shape enables closed-form solutions to many problems in applied probability. However, the classical definition of high-dimensional FGM copula does not enable a straightforward understanding of the effect of the copula parameters on the dependence, nor a geometric understanding of their admissible range. We circumvent this issue by studying the FGM copula from a probabilistic approach based on multivariate Bernoulli distributions. This paper studies high-dimensional exchangeable FGM copulas, a subclass of FGM copulas. We show that dependence parameters of exchangeable FGM can be expressed as convex hulls of a finite number of extreme points and establish partial orders for different exchangeable FGM copulas (including maximal and minimal dependence). We also leverage the probabilistic interpretation to develop efficient sampling and estimating procedures and provide a simulation study. Throughout, we discover geometric interpretations of the copula parameters that assist one in decoding the dependence of high-dimensional exchangeable FGM copulas.

math.ST

Geographic ratemaking with spatial embeddings

Spatial data is a rich source of information for actuarial applications: knowledge of a risk's location could improve an insurance company's ratemaking, reserving or risk management processes. Insurance companies with high exposures in a territory typically have a competitive advantage since they may use historical losses in a region to model spatial risk non-parametrically. Relying on geographic losses is problematic for areas where past loss data is unavailable. This paper presents a method based on data (instead of smoothing historical insurance claim losses) to construct a geographic ratemaking model. In particular, we construct spatial features within a complex representation model, then use the features as inputs to a simpler predictive model (like a generalized linear model). Our approach generates predictions with smaller bias and smaller variance than other spatial interpolation models such as bivariate splines in most situations. This method also enables us to generate rates in territories with no historical experience.

stat.AP

Rethinking Representations in P&C Actuarial Science with Deep Neural Networks

Insurance companies gather a growing variety of data for use in the insurance process, but most traditional ratemaking models are not designed to support them. In particular, many emerging data sources (text, images, sensors) may complement traditional data to provide better insights to predict the future losses in an insurance contract. This paper presents some of these emerging data sources and presents a unified framework for actuaries to incorporate these in existing ratemaking models. Our approach stems from representation learning, whose goal is to create representations of raw data. A useful representation will transform the original data into a dense vector space where the ultimate predictive task is simpler to model. Our paper presents methods to transform non-vectorial data into vectorial representations and provides examples for actuarial science.

stat.AP