SearcharxivSearch

arXiv subjects

Philippe Naveau

Publications and source records attributed to Philippe Naveau.

At least 19 recordsLinked to original sources

Bayesian spatial modelling framework for assessing residential flood risk in property insurance

Spatial heterogeneity in insurance risk modelling is often represented using coarse areal structures, which can obscure fine-scale patterns critical for accurate risk assessment. This study introduces a point-referenced Bayesian framework to model claim occurrence and severity at the policyholder level, avoiding reliance on predefined geographic aggregation. Drawing on a large French insurance portfolio combined with high-resolution environmental variables, rainfall records, and institutional hazard maps, we compare a benchmark GLM with several discrete Bayesian specifications, including independent random effects, intrinsic conditional autoregressive (iCAR) and Besag-York-Mollie (BYM) models, and a continuously indexed Gaussian random field constructed using the stochastic partial differential equation (SPDE) approach. Inference is performed using Integrated Nested Laplace Approximation (INLA), enabling efficient estimation of latent spatial fields and non-linear covariate effects. Our results show that accounting for spatial dependence substantially improves occurrence modelling, while gains in severity prediction are more limited. The SPDE formulation further outperforms areal models by capturing sub-municipal risk gradients and reducing artefacts induced by arbitrary geographic partitioning. By conditioning on detailed building-level attributes, we isolate the contribution of latent spatial effects, refine the interpretation of observed covariates, and improve the allocation of risk premiums across the portfolio. In addition to enhanced predictive performance, the framework provides coherent uncertainty quantification and supports tail-risk assessment. To our knowledge, this is the first application of point-referenced SPDE models to flood insurance, offering a scalable statistical alternative for pricing and managing risks with strong spatial structure.

stat.AP

Benchmark of Likelihood-Free Inference Methods based on Neural and Optimal Transport Approaches

Simulation-based inference (SBI) has become an increasingly important framework for parameter estimation in models for which simulation is feasible, including cases where likelihood evaluation is unavailable or costly. While recent work has introduced benchmark frameworks to compare likelihood-free methods, these studies often do not account for structural features such as heavy-tails or discreteness. In this article, we investigate how the performance of likelihood-free inference methods depends on these structural properties. We consider four approaches: MLE, NBE, EOT and AW--NBE and evaluate them using simulations. This study highlights the importance of carefully selecting evaluation tools in the presence of extremes and discrete data.

stat.ME

Likelihood-Free Inference for Multivariate Generalized Pareto Models

Likelihood-based inference for multivariate extreme-value models is often unreliable or infeasible when likelihoods are intractable or supports are discrete. This challenge is particularly acute for multivariate discrete generalized Pareto models, where both marginal tail behavior and dependence must be inferred from sparse exceedance samples. We propose a two-stage likelihood-free inference procedure, termed AW--NBE (Adaptive Wasserstein Neural Bayes Estimator), that combines neural Bayes estimation with a targeted optimal transport refinement step based on the Sinkhorn discrepancy. In the first stage, a neural Bayes estimator trained on simulated data provides fast and stable initial parameter estimates. In the second stage, these estimates are locally refined by minimizing the Sinkhorn divergence between the empirical distributions of observed and simulated exceedances. This refinement reduces the Sinkhorn discrepancy between the empirical distributions of observed and simulated exceedances, while preserving dependence features learned by the neural estimator. Model adequacy is assessed using new optimal transport based multivariate Q--Q and potential diagnostics. Applications to financial log-returns and Swiss dry spell exceedances suggest that AW--NBE can improve parameter inferences compared to estimation using solely, either the Sinkhorn discrepancy, or the standard neural Bayes estimators and censored likelihood estimation.

stat.AP

A duration-augmented binary Markov chain for rainfall occurrence with long dry spells

Simulating realistic wet and dry spells is central in weather generators and climate-impact studies. While finite-order Markov chains are standard, they often fail to reproduce persistent dry conditions due to their inherent subexponential decay. We model rainfall occurrence by introducing a duration-augmented binary Markov chain. We establish a link with alternating renewal chains, enabling flexible parametric modelling of wet and dry spell duration distribution. We model those using two regime-adapted specifications from the general class of extended Generalized Pareto Distributions, yielding flexible tail behaviour across various climates. We use estimation methods adapted to each specification. Our model is applied to around 200 stations in the South of Europe spanning diverse Mediterranean and continental climates. We compare this framework to standard Markov models in characterising persistence and high-quantile extrapolation. The approach is generic, extending naturally to multi-state settings or other binary sequence applications in environmental statistics.

stat.ME

A sub-asymptotic model for bivariate threshold exceedances

Extreme value theory offers a statistical framework for quantifying the risk of rare events, with the generalized Pareto (GP) distribution providing the canonical limit model for univariate threshold exceedances. In many applications, however, extremes are intrinsically multivariate, requiring models that capture both marginal tail behaviours and joint extremal dependencies. Under asymptotic dependence, the multivariate GP distribution represents a suitable modelling family, but when asymptotic independence arises, sub-asymptotic models are needed. In this work, we propose and study a flexible sub-asymptotic parametric class to model bivariate threshold exceedances. Our new model accommodates a broad range of tail dependence behaviours and contains the standardised multivariate GP distribution as a limiting case while retaining margins that converge to univariate GP tails. Our formulation allows extremal dependence to evolve naturally with the marginal parameters on the original data scale, facilitating direct computation and interpretation of failure probabilities. Model inference is done via a likelihood-free neural Bayes estimation approach, with tailored prior specifications. An extensive simulation study and an application to Belgian rainfall extremes illustrate the estimation framework and the flexibility of the model.

stat.ME

Contributions of geolocated weather and building related data for insurance assessment of flood risks

Floods rank among the costliest natural hazards, causing over USD 100 billion in insured losses between 2013 and 2023. In France, persistent deficits in the natural catastrophe scheme highlight the need for accurate, building-scale flood risk assessment. Insurers typically rely on frequency-severity models supported by hazard maps and regional climate indicators. However, previous studies show that such large-scale variables explain only a limited share of the variability in individual flood losses. This study evaluates the marginal contribution of multiple georeferenced data layers to modeling flood claim occurrence and severity in a large French home insurance portfolio. Starting from a baseline model based on standard underwriting information, we sequentially introduce climate-expert variables, extreme rainfall indicators, and fine-scale geolocated building and environmental attributes. The analysis focuses on a practical setting in which insurers cannot deploy full hydrological or hydraulic catastrophe models because of budgetary, licensing, or operational constraints. Results show that rainfall-based indicators, particularly a newly constructed metric capturing intense local precipitation, substantially improve claim modeling performance. Building and environmental variables further enhance occurrence prediction. Overall, the findings demonstrate how high-resolution geolocated data improve exposure and vulnerability assessment, complement official flood maps, and provide insurers with an operational framework for refining flood risk evaluation and pricing.

stat.AP

A Kullback-Leibler divergence test for multivariate extremes: theory and practice

Testing whether two multivariate samples exhibit the same extremal behavior is an important problem in various fields including environmental and climate sciences. While several ad-hoc approaches exist in the literature, they often lack theoretical justification and statistical guarantees. On the other hand, extreme value theory provides the theoretical foundation for constructing asymptotically justified tests. We combine this theory with Kullback-Leibler divergence, a fundamental concept in information theory and statistics, to propose a test for equality of extremal dependence structures in practically relevant directions. Under suitable assumptions, we derive the limiting distributions of the proposed statistic under null and alternative hypotheses. Importantly, our test is fast to compute and easy to interpret by practitioners, making it attractive in applications. Simulations provide evidence of the power of our test. In a case study, we apply our method to show the strong impact of seasons on the strength of dependence between different aggregation periods (daily versus hourly) of heavy rainfall in France.

math.ST

A parsimonious tail compliant multiscale statistical model for aggregated rainfall

Modeling rainfall intensity distributions across aggregation scales (from sub-hourly to weekly) is essential for hydrological risk analysis and IDF curves. Aggregation naturally imposes mathematical constraints: return levels must be ordered by time scale, as daily accumulations necessarily exceed sub-daily ones. From a statistical perspective, each aggregation step should ideally not require additional parameters, yet parsimonious models describing the full distribution remain scarce, as most literature focuses on seasonal block maxima. In this study, we propose a parsimonious framework to model all rainfall intensities (low to large) across scales. We utilize the Extended Generalized Pareto Distribution (EGPD), which aligns with extreme value theory for both tails while remaining flexible for the bulk of the distribution. We establish a general result on the behavior of EGPD variables under various aggregation procedures. To overcome the difficulty of direct likelihood inference, we link the EGPD class to Poisson compound sums. This allows the use of the Panjer algorithm for efficient composite likelihood evaluation. Our approach ensures that return levels do not cross across scales and enables estimation for return periods below annual or seasonal levels. We demonstrate the method using sub-hourly series from six French stations with diverse climates. Only eight parameters are needed per station to capture scales from six minutes to three days. IDF curves above and below the annual scale are provided.

stat.AP

Multivariate distributional modeling of low, moderate, and large intensities without threshold selection steps

In fields such as hydrology and climatology, modelling the entire distribution of positive data is essential, as stakeholders require insights into the full range of values, from low to extreme. Traditional approaches often segment the distribution into separate regions, which introduces subjectivity and limits coherence. This is especially true when dealing with multivariate data. In line with multivariate extreme value theory, this paper presents a unified, threshold-free framework for modelling marginal behaviours and dependence structures based on an extended generalized Pareto distribution (EGPD). We propose decomposing multivariate data into radial and angular components. The radial component is modelled using a semi-parametric EGPD and the angular distribution is permitted to vary conditionally. This approach allows for sufficiently flexible dependence modelling. The hierarchical structure of the model facilitates the inference process. First, we combine classical maximum likelihood estimation (MLE) methods with semi-parametric approaches based on Bernstein polynomials to estimate the distribution of the radial component. Then, we use multivariate regression techniques to estimate the angular component's parameters. The model is evaluated through synthetic simulations and applied to hydrological datasets to exemplify its capacity to capture heavy-tailed marginals and complex multivariate dependencies without threshold specification.

stat.ME

Joint modeling of low and high extremes using a multivariate extended generalized Pareto distribution

In most risk assessment studies, it is important to accurately capture the entire distribution of the multivariate random vector of interest from low to high values. For example, in climate sciences, low precipitation events may lead to droughts, while heavy rainfall may generate large floods, and both of these extreme scenarios can have major impacts on the safety of people and infrastructure, as well as agricultural or other economic sectors. In the univariate case, the extended generalized Pareto distribution (eGPD) was specifically developed to accurately model low, moderate, and high precipitation intensities, while bypassing the threshold selection procedure usually conducted in extreme-value analyses. In this work, we extend this approach to the multivariate case. The proposed multivariate eGPD has the following appealing properties: (1) its marginal distributions behave like univariate eGPDs; (2) its lower and upper joint tails comply with multivariate extreme-value theory, with key parameters separately controlling dependence in each joint tail; and (3) the model allows for fast simulation and is thus amenable to simulation-based inference. We propose estimating model parameters by leveraging modern neural approaches, where a neural network, once trained, can provide point estimates, credible intervals, or full posterior approximations in a fraction of a second. Our new methodology is illustrated by application to daily rainfall times series data from the Netherlands. The proposed model is shown to provide satisfactory marginal and dependence fits from low to high quantiles.

stat.ME

Multivariate Discrete Generalized Pareto Distributions: Theory, Simulation, and Applications to Dry spells

This article extends the multivariate extreme value theory (MEVT) to discrete settings, focusing on the generalized Pareto distribution (GPD) as a foundational tool. The purpose of the study is to enhance the understanding of extreme discrete count data representation, particularly for discrete exceedances over thresholds, defining and using multivariate discrete Pareto distributions (MDGPD). Through theoretical results and illustrative examples, we outline the construction and properties of MDGPDs, providing practical insights into simulation techniques and data fitting approaches using recent likelihood-free inference methods. This framework broadens the toolkit for modeling extreme events, offering robust methodologies for analyzing multivariate discrete data with extreme values. To illustrate its practical relevance, we present an application of this method to drought analysis, addressing a growing concern in Europe.

stat.ME

Multi-site modelling and reconstruction of past extreme skew surges along the French Atlantic coast

Appropriate modelling of extreme skew surges is crucial, particularly for coastal risk management. Our study focuses on modelling extreme skew surges along the French Atlantic coast, with a particular emphasis on investigating the extremal dependence structure between stations. We employ the peak-over-threshold framework, where a multivariate extreme event is defined whenever at least one location records a large value, though not necessarily all stations simultaneously. A novel method for determining an appropriate level (threshold) above which observations can be classified as extreme is proposed. Two complementary approaches are explored. First, the multivariate generalized Pareto distribution is employed to model extremes, leveraging its properties to derive a generative model that predicts extreme skew surges at one station based on observed extremes at nearby stations. Second, a novel extreme regression framework is assessed for point predictions. This specific regression framework enables accurate point predictions using only the 'angle' of input variables, i.e., input variables divided by their norms. The ultimate objective is to reconstruct historical skew surge time series at stations with limited data. This is achieved by integrating extreme skew surge data from stations with longer records, such as Brest and Saint-Nazaire, which provide over 150 years of observations.

stat.AP

Multivariate extreme value theory

When passing from the univariate to the multivariate setting, modelling extremes becomes much more intricate. In this introductory exposition, classical multivariate extreme value theory is presented from the point of view of multivariate excesses over high thresholds as modelled by the family of multivariate generalized Pareto distributions. The formulation in terms of failure sets in the sample space intersecting the sample cloud leads to the over-arching perspective of point processes. Max-stable or generalized extreme value distributions are finally obtained as limits of vectors of componentwise maxima by considering the event that a certain region of the sample space does not contain any observation.

math.ST

MERCURY: A fast and versatile multi-resolution based global emulator of compound climate hazards

High-impact climate damages are often driven by compounding climate conditions. For example, elevated heat stress conditions can arise from a combination of high humidity and temperature. To explore future changes in compounding hazards under a range of climate scenarios and with large ensembles, climate emulators can provide light-weight, data-driven complements to Earth System Models. Yet, only a few existing emulators can jointly emulate multiple climate variables. In this study, we present the Multi-resolution EmulatoR for CompoUnd climate Risk analYsis: MERCURY. MERCURY extends multi-resolution analysis to a spatio-temporal framework for versatile emulation of multiple variables. MERCURY leverages data-driven, image compression techniques to generate emulations in a memory-efficient manner. MERCURY consists of a regional component that represents the monthly, regional response of a given variable to yearly Global Mean Temperature (GMT) using a probabilistic regression based additive model, resolving regional cross-correlations. It then adapts a reverse lifting-scheme operator to jointly spatially disaggregate regional, monthly values to grid-cell level. We demonstrate MERCURY's capabilities on representing the humid-heat metric, Wet Bulb Globe Temperature, as derived from temperature and relative humidity emulations. The emulated WBGT spatial correlations correspond well to those of ESMs and the 95% and 97.5% quantiles of WBGT distributions are well captured, with an average of 5% deviation. MERCURY's setup allows for region-specific emulations from which one can efficiently "zoom" into the grid-cell level across multiple variables by means of the reverse lifting-scheme operator. This circumvents the traditional problem of having to emulate complete, global-fields of climate data and resulting storage requirements.

physics.ao-ph

Distributional Regression U-Nets for the Postprocessing of Precipitation Ensemble Forecasts

Accurate precipitation forecasts have a high socio-economic value due to their role in decision-making in various fields such as transport networks and farming. We propose a global statistical postprocessing method for grid-based precipitation ensemble forecasts. This U-Net-based distributional regression method predicts marginal distributions in the form of parametric distributions inferred by scoring rule minimization. Distributional regression U-Nets are compared to state-of-the-art postprocessing methods for daily 21-h forecasts of 3-h accumulated precipitation over the South of France. Training data comes from the M\'et\'eo-France weather model AROME-EPS and spans 3 years. A practical challenge appears when consistent data or reforecasts are not available. Distributional regression U-Nets compete favorably with the raw ensemble. In terms of continuous ranked probability score, they reach a performance comparable to quantile regression forests (QRF). However, they are unable to provide calibrated forecasts in areas associated with high climatological precipitation. In terms of predictive power for heavy precipitation events, they outperform both QRF and semi-parametric QRF with tail extensions.

stat.ML

Proper Scoring Rules for Multivariate Probabilistic Forecasts based on Aggregation and Transformation

Proper scoring rules are an essential tool to assess the predictive performance of probabilistic forecasts. However, propriety alone does not ensure an informative characterization of predictive performance and it is recommended to compare forecasts using multiple scoring rules. With that in mind, interpretable scoring rules providing complementary information are necessary. We formalize a framework based on aggregation and transformation to build interpretable multivariate proper scoring rules. Aggregation-and-transformation-based scoring rules are able to target specific features of the probabilistic forecasts; which improves the characterization of the predictive performance. This framework is illustrated through examples taken from the literature and studied using numerical experiments showcasing its benefits. In particular, it is shown that it can help bridge the gap between proper scoring rules and spatial verification tools.

stat.ME

Causal Discovery in Multivariate Extremes with a Hydrological Analysis of Swiss River Discharges

Causal asymmetry is based on the principle that an event is a cause only if its absence would not have been a cause. From there, uncovering causal effects becomes a matter of comparing a well-defined score in both directions. Motivated by studying causal effects at extreme levels of a multivariate random vector, we propose to construct a model-agnostic causal score relying solely on the assumption of the existence of a max-domain of attraction. Based on a representation of a Generalized Pareto random vector, we construct the causal score as the Wasserstein distance between the margins and a well-specified random variable. The proposed methodology is illustrated on a hydrologically simulated dataset of different characteristics of catchments in Switzerland: discharge, precipitation, and snowmelt.

stat.ME

Robust and non asymptotic estimation of probability weighted moments with application to extreme value analysis

In extreme value theory and other related risk analysis fields, probability weighted moments (PWM) have been frequently used to estimate the parameters of classical extreme value distributions. This method-of-moment technique can be applied when second moments are finite, a reasonable assumption in many environmental domains like climatological and hydrological studies. Three advantages of PWM estimators can be put forward: their simple interpretations, their rapid numerical implementation and their close connection to the well-studied class of U-statistics. Concerning the later, this connection leads to precise asymptotic properties, but non asymptotic bounds have been lacking when off-the-shelf techniques (Chernoff method) cannot be applied, as exponential moment assumptions become unrealistic in many extreme value settings. In addition, large values analysis is not immune to the undesirable effect of outliers, for example, defective readings in satellite measurements or possible anomalies in climate model runs. Recently, the treatment of outliers has sparked some interest in extreme value theory, but results about finite sample bounds in a robust extreme value theory context are yet to be found, in particular for PWMs or tail index estimators. In this work, we propose a new class of robust PWM estimators, inspired by the median-of-means framework of Devroye et al. (2016). This class of robust estimators is shown to satisfy a sub-Gaussian inequality when the assumption of finite second moments holds. Such non asymptotic bounds are also derived under the general contamination model. Our main proposition confirms theoretically a trade-off between efficiency and robustness. Our simulation study indicates that, while classical estimators of PWMs can be highly sensitive to outliers.

math.ST