SearcharxivSearch

arXiv subjects

Juliette Legrand

Publications and source records attributed to Juliette Legrand.

6 recordsLinked to original sources

Extrapolation of extreme covariates in generalized additive regression using extreme-value theory

We propose methods to enhance the predictive performance of generalized additive models (GAMs) in the context of covariate extrapolation, where predictions rely on covariates beyond their observed range. When using predictive models such as GAMs, shifts in the covariate distribution between training and prediction datasets can occur. Ignoring this issue may lead to inaccurate predictions in the tail of the covariate distributions. For example, this problem is particularly critical in climate-change scenarios, where covariates simulated from future climate scenarios are likely to contain more extreme conditions. Our approach integrates GAMs for the bulk of covariate distributions with asymptotic models from multivariate extreme-value theory at high covariate values. We consider binary responses based on a latent variable assumption, and also continuous responses. For large values of the covariates, on a specific marginal scale motivated by extreme-value theory the latent variable or continuous response is assumed to depend linearly on the covariates with an additive error term, when using an appropriate link function. In an application to wildfires in Europe, we explore how the new method can improve predictions, using environmental and meteorological covariates.

stat.ME

Bayesian spatial modelling framework for assessing residential flood risk in property insurance

Spatial heterogeneity in insurance risk modelling is often represented using coarse areal structures, which can obscure fine-scale patterns critical for accurate risk assessment. This study introduces a point-referenced Bayesian framework to model claim occurrence and severity at the policyholder level, avoiding reliance on predefined geographic aggregation. Drawing on a large French insurance portfolio combined with high-resolution environmental variables, rainfall records, and institutional hazard maps, we compare a benchmark GLM with several discrete Bayesian specifications, including independent random effects, intrinsic conditional autoregressive (iCAR) and Besag-York-Mollie (BYM) models, and a continuously indexed Gaussian random field constructed using the stochastic partial differential equation (SPDE) approach. Inference is performed using Integrated Nested Laplace Approximation (INLA), enabling efficient estimation of latent spatial fields and non-linear covariate effects. Our results show that accounting for spatial dependence substantially improves occurrence modelling, while gains in severity prediction are more limited. The SPDE formulation further outperforms areal models by capturing sub-municipal risk gradients and reducing artefacts induced by arbitrary geographic partitioning. By conditioning on detailed building-level attributes, we isolate the contribution of latent spatial effects, refine the interpretation of observed covariates, and improve the allocation of risk premiums across the portfolio. In addition to enhanced predictive performance, the framework provides coherent uncertainty quantification and supports tail-risk assessment. To our knowledge, this is the first application of point-referenced SPDE models to flood insurance, offering a scalable statistical alternative for pricing and managing risks with strong spatial structure.

stat.AP

Contributions of geolocated weather and building related data for insurance assessment of flood risks

Floods rank among the costliest natural hazards, causing over USD 100 billion in insured losses between 2013 and 2023. In France, persistent deficits in the natural catastrophe scheme highlight the need for accurate, building-scale flood risk assessment. Insurers typically rely on frequency-severity models supported by hazard maps and regional climate indicators. However, previous studies show that such large-scale variables explain only a limited share of the variability in individual flood losses. This study evaluates the marginal contribution of multiple georeferenced data layers to modeling flood claim occurrence and severity in a large French home insurance portfolio. Starting from a baseline model based on standard underwriting information, we sequentially introduce climate-expert variables, extreme rainfall indicators, and fine-scale geolocated building and environmental attributes. The analysis focuses on a practical setting in which insurers cannot deploy full hydrological or hydraulic catastrophe models because of budgetary, licensing, or operational constraints. Results show that rainfall-based indicators, particularly a newly constructed metric capturing intense local precipitation, substantially improve claim modeling performance. Building and environmental variables further enhance occurrence prediction. Overall, the findings demonstrate how high-resolution geolocated data improve exposure and vulnerability assessment, complement official flood maps, and provide insurers with an operational framework for refining flood risk evaluation and pricing.

stat.AP

Pareto processes for threshold exceedances in spatial extremes

We review some recent development in the theory of spatial extremes related to Pareto Processes and modeling of threshold exceedances. We provide theoretical background, methodology for modeling, simulation and inference as well as an illustration to wave height modelling. This preprint is an author version of a chapter to appear in a collaborative book.

math.ST

Assessing Extreme Risk using Stochastic Simulation of Extremes

Risk management is particularly concerned with extreme events, but analysing these events is often hindered by the scarcity of data, especially in a multivariate context. This data scarcity complicates risk management efforts. Various tools can assess the risk posed by extreme events, even under extraordinary circumstances. This paper studies the evaluation of univariate risk for a given risk factor using metrics that account for its asymptotic dependence on other risk factors. Data availability is crucial, particularly for extreme events where it is often limited by the nature of the phenomenon itself, making estimation challenging. To address this issue, two non-parametric simulation algorithms based on multivariate extreme theory are developed. These algorithms aim to extend a sample of extremes jointly and conditionally for asymptotically dependent variables using stochastic simulation and multivariate Generalised Pareto Distributions. The approach is illustrated with numerical analyses of both simulated and real data to assess the accuracy of extreme risk metric estimations.

stat.ME

Evaluation of binary classifiers for asymptotically dependent and independent extremes

Machine learning classification methods usually assume that all possible classes are sufficiently present within the training set. Due to their inherent rarities, extreme events are always under-represented and classifiers tailored for predicting extremes need to be carefully designed to handle this under-representation. In this paper, we address the question of how to assess and compare classifiers with respect to their capacity to capture extreme occurrences. This is also related to the topic of scoring rules used in forecasting literature. In this context, we propose and study a risk function adapted to extremal classifiers. The inferential properties of our empirical risk estimator are derived under the framework of multivariate regular variation and hidden regular variation. A simulation study compares different classifiers and indicates their performance with respect to our risk function. To conclude, we apply our framework to the analysis of extreme river discharges in the Danube river basin. The application compares different predictive algorithms and test their capacity at forecasting river discharges from other river stations.

stat.ME