SearcharxivSearch

arXiv subjects

Pierre Ailliot

Publications and source records attributed to Pierre Ailliot.

16 recordsLinked to original sources

Bayesian spatial modelling framework for assessing residential flood risk in property insurance

Spatial heterogeneity in insurance risk modelling is often represented using coarse areal structures, which can obscure fine-scale patterns critical for accurate risk assessment. This study introduces a point-referenced Bayesian framework to model claim occurrence and severity at the policyholder level, avoiding reliance on predefined geographic aggregation. Drawing on a large French insurance portfolio combined with high-resolution environmental variables, rainfall records, and institutional hazard maps, we compare a benchmark GLM with several discrete Bayesian specifications, including independent random effects, intrinsic conditional autoregressive (iCAR) and Besag-York-Mollie (BYM) models, and a continuously indexed Gaussian random field constructed using the stochastic partial differential equation (SPDE) approach. Inference is performed using Integrated Nested Laplace Approximation (INLA), enabling efficient estimation of latent spatial fields and non-linear covariate effects. Our results show that accounting for spatial dependence substantially improves occurrence modelling, while gains in severity prediction are more limited. The SPDE formulation further outperforms areal models by capturing sub-municipal risk gradients and reducing artefacts induced by arbitrary geographic partitioning. By conditioning on detailed building-level attributes, we isolate the contribution of latent spatial effects, refine the interpretation of observed covariates, and improve the allocation of risk premiums across the portfolio. In addition to enhanced predictive performance, the framework provides coherent uncertainty quantification and supports tail-risk assessment. To our knowledge, this is the first application of point-referenced SPDE models to flood insurance, offering a scalable statistical alternative for pricing and managing risks with strong spatial structure.

stat.AP

Contributions of geolocated weather and building related data for insurance assessment of flood risks

Floods rank among the costliest natural hazards, causing over USD 100 billion in insured losses between 2013 and 2023. In France, persistent deficits in the natural catastrophe scheme highlight the need for accurate, building-scale flood risk assessment. Insurers typically rely on frequency-severity models supported by hazard maps and regional climate indicators. However, previous studies show that such large-scale variables explain only a limited share of the variability in individual flood losses. This study evaluates the marginal contribution of multiple georeferenced data layers to modeling flood claim occurrence and severity in a large French home insurance portfolio. Starting from a baseline model based on standard underwriting information, we sequentially introduce climate-expert variables, extreme rainfall indicators, and fine-scale geolocated building and environmental attributes. The analysis focuses on a practical setting in which insurers cannot deploy full hydrological or hydraulic catastrophe models because of budgetary, licensing, or operational constraints. Results show that rainfall-based indicators, particularly a newly constructed metric capturing intense local precipitation, substantially improve claim modeling performance. Building and environmental variables further enhance occurrence prediction. Overall, the findings demonstrate how high-resolution geolocated data improve exposure and vulnerability assessment, complement official flood maps, and provide insurers with an operational framework for refining flood risk evaluation and pricing.

stat.AP

A parsimonious tail compliant multiscale statistical model for aggregated rainfall

Modeling rainfall intensity distributions across aggregation scales (from sub-hourly to weekly) is essential for hydrological risk analysis and IDF curves. Aggregation naturally imposes mathematical constraints: return levels must be ordered by time scale, as daily accumulations necessarily exceed sub-daily ones. From a statistical perspective, each aggregation step should ideally not require additional parameters, yet parsimonious models describing the full distribution remain scarce, as most literature focuses on seasonal block maxima. In this study, we propose a parsimonious framework to model all rainfall intensities (low to large) across scales. We utilize the Extended Generalized Pareto Distribution (EGPD), which aligns with extreme value theory for both tails while remaining flexible for the bulk of the distribution. We establish a general result on the behavior of EGPD variables under various aggregation procedures. To overcome the difficulty of direct likelihood inference, we link the EGPD class to Poisson compound sums. This allows the use of the Panjer algorithm for efficient composite likelihood evaluation. Our approach ensures that return levels do not cross across scales and enables estimation for return periods below annual or seasonal levels. We demonstrate the method using sub-hourly series from six French stations with diverse climates. Only eight parameters are needed per station to capture scales from six minutes to three days. IDF curves above and below the annual scale are provided.

stat.AP

Note on a non-parametric method for change-point detection

The purpose of this note is to present in details R codes to implement a non-parametric method for change-point detection. The proposed approach is validated from various perspectives using simulations. This method is a competitor to that of Pettitt ([3]) and is, like the latter, based on the Wilcoxon-Mann-Whitney test. It is used in [4] for the study of relatively short time series obtained from measurements on cores sampled in the bay of Brest.

stat.AP

EM algorithm for generalized Ridge regression with spatial covariates

The generalized Ridge penalty is a powerful tool for dealing with overfitting and for high-dimensional regressions. The generalized Ridge regression can be derived as the mean of a posterior distribution with a Normal prior and a given covariance matrix. The covariance matrix controls the structure of the coefficients, which depends on the particular application. For example, it is appropriate to assume that the coefficients have a spatial structure in spatial applications. This study proposes an expectation-maximization algorithm for estimating generalized Ridge parameters whose covariance structure depends on specific parameters. We focus on three cases: diagonal (when the covariance matrix is diagonal with constant elements), Matérn, and conditional autoregressive covariances. A simulation study is conducted to evaluate the performance of the proposed method, and then the method is applied to predict ocean wave heights using wind conditions.

stat.ME

Learning the spatio-temporal relationship between wind and significant wave height using deep learning

Ocean wave climate has a significant impact on near-shore and off-shore human activities, and its characterisation can help in the design of ocean structures such as wave energy converters and sea dikes. Therefore, engineers need long time series of ocean wave parameters. Numerical models are a valuable source of ocean wave data; however, they are computationally expensive. Consequently, statistical and data-driven approaches have gained increasing interest in recent decades. This work investigates the spatio-temporal relationship between North Atlantic wind and significant wave height (Hs) at an off-shore location in the Bay of Biscay, using a two-stage deep learning model. The first step uses convolutional neural networks (CNNs) to extract the spatial features that contribute to Hs. Then, long short-term memory (LSTM) is used to learn the long-term temporal dependencies between wind and waves.

stat.ML

An algorithm for non-parametric estimation in state-space models

State-space models are ubiquitous in the statistical literature since they provide a flexible and interpretable framework for analyzing many time series. In most practical applications, the state-space model is specified through a parametric model. However, the specification of such a parametric model may require an important modeling effort or may lead to models which are not flexible enough to reproduce all the complexity of the phenomenon of interest. In such situations, an appealing alternative consists in inferring the state-space model directly from the data using a non-parametric framework. The recent developments of powerful simulation techniques have permitted to improve the statistical inference for parametric state-space models. It is proposed to combine two of these techniques, namely the Stochastic Expectation-Maximization (SEM) algorithm and Sequential Monte Carlo (SMC) approaches, for non-parametric estimation in state-space models. The performance of the proposed algorithm is assessed through simulations on toy models and an application to environmental data is discussed.

stat.ME

A Review of Innovation-Based Methods to Jointly Estimate Model and Observation Error Covariance Matrices in Ensemble Data Assimilation

Data assimilation combines forecasts from a numerical model with observations. Most of the current data assimilation algorithms consider the model and observation error terms as additive Gaussian noise, specified by their covariance matrices Q and R, respectively. These error covariances, and specifically their respective amplitudes, determine the weights given to the background (i.e., the model forecasts) and to the observations in the solution of data assimilation algorithms (i.e., the analysis). Consequently, Q and R matrices significantly impact the accuracy of the analysis. This review aims to present and to discuss, with a unified framework, different methods to jointly estimate the Q and R matrices using ensemble-based data assimilation techniques. Most of the methodologies developed to date use the innovations, defined as differences between the observations and the projection of the forecasts onto the observation space. These methodologies are based on two main statistical criteria: (i) the method of moments, in which the theoretical and empirical moments of the innovations are assumed to be equal, and (ii) methods that use the likelihood of the observations, themselves contained in the innovations. The reviewed methods assume that innovations are Gaussian random variables, although extension to other distributions is possible for likelihood-based methods. The methods also show some differences in terms of levels of complexity and applicability to high-dimensional systems. The conclusion of the review discusses the key challenges to further develop estimation methods for Q and R. These challenges include taking into account time-varying error covariances, using limited observational coverage, estimating additional deterministic error terms, or accounting for correlated noises.

stat.ME

An efficient particle-based method for maximum likelihood estimation in nonlinear state-space models

Data assimilation methods aim at estimating the state of a system by combining observations with a physical model. When sequential data assimilation is considered, the joint distribution of the latent state and the observations is described mathematically using a state-space model, and filtering or smoothing algorithms are used to approximate the conditional distribution of the state given the observations. The most popular algorithms in the data assimilation community are based on the Ensemble Kalman Filter and Smoother (EnKF/EnKS) and its extensions. In this paper we investigate an alternative approach where a Conditional Particle Filter (CPF) is combined with Backward Simulation (BS). This allows to explore efficiently the latent space and simulate quickly relevant trajectories of the state conditionally to the observations. We also tackle the difficult problem of parameter estimation. Indeed, the models generally involve statistical parameters in the physical models and/or in the stochastic models for the errors. These parameters strongly impact the results of the data assimilation algorithm and there is a need for an efficient method to estimate them. Expectation-Maximization (EM) is the most classical algorithm in the statistical literature to estimate the parameters in models with latent variables. It consists in updating sequentially the parameters by maximizing a likelihood function where the state is approximated using a smoothing algorithm. In this paper, we propose an original Stochastic Expectation-Maximization (SEM) algorithm combined to the CPF-BS smoother to estimate the statistical parameters. We show on several toy models that this algorithm provides, with reasonable computational cost, accurate estimations of the statistical parameters and the state in highly nonlinear state-space models, where the application of EM algorithms using EnKS is limited. We also provide a Python source code of the algorithm.

stat.ME

Locally-adapted convolution-based super-resolution of irregularly-sampled ocean remote sensing data

Super-resolution is a classical problem in image processing, with numerous applications to remote sensing image enhancement. Here, we address the super-resolution of irregularly-sampled remote sensing images. Using an optimal interpolation as the low-resolution reconstruction, we explore locally-adapted multimodal convolutional models and investigate different dictionary-based decompositions, namely based on principal component analysis (PCA), sparse priors and non-negativity constraints. We consider an application to the reconstruction of sea surface height (SSH) fields from two information sources, along-track altimeter data and sea surface temperature (SST) data. The reported experiments demonstrate the relevance of the proposed model, especially locally-adapted parametrizations with non-negativity constraints, to outperform optimally-interpolated reconstructions.

stat.ML

Dependent time changed processes with applications to nonlinear ocean waves

Many records in environmental sciences exhibit asymmetric trajectories and there is a need for simple and tractable models which can reproduce such features. In this paper we explore an approach based on applying both a time change and a marginal transformation on Gaussian processes. The main originality of the proposed model is that the time change depends on the observed trajectory. We first show that the proposed model is stationary and ergodic and provide an explicit characterization of the stationary distribution. This result is then used to build both parametric and non-parametric estimates of the time change function whereas the estimation of the marginal transformation is based on up-crossings. Simulation results are provided to assess the quality of the estimates. The model is applied to shallow water wave data and it is shown that the fitted model is able to reproduce important statistics of the data such as its spectrum and marginal distribution which are important quantities for practical applications. An important benefit of the proposed model is its ability to reproduce the observed asymmetries between the crest and the troughs and between the front and the back of the waves by accelerating the chronometer in the crests and in the front of the waves.

stat.ME

Consistency of the maximum likelihood estimate for Non-homogeneous Markov-switching models

Many nonlinear time series models have been proposed in the last decades. Among them, the models with regime switchings provide a class of versatile and interpretable models which have received a particular attention in the literature. In this paper, we consider a large family of such models which generalize the well known Markov-switching AutoRegressive (MS-AR) by allowing non-homogeneous switching and encompass Threshold AutoRegressive (TAR) models. We prove various theoretical results related to the stability of these models and the asymptotic properties of the Maximum Likelihood Estimates (MLE). The ability of the model to catch complex nonlinearities is then illustrated on various time series.

stat.AP

Modeling extreme values of processes observed at irregular time steps: Application to significant wave height

This work is motivated by the analysis of the extremal behavior of buoy and satellite data describing wave conditions in the North Atlantic Ocean. The available data sets consist of time series of significant wave height (Hs) with irregular time sampling. In such a situation, the usual statistical methods for analyzing extreme values cannot be used directly. The method proposed in this paper is an extension of the peaks over threshold (POT) method, where the distribution of a process above a high threshold is approximated by a max-stable process whose parameters are estimated by maximizing a composite likelihood function. The efficiency of the proposed method is assessed on an extensive set of simulated data. It is shown, in particular, that the method is able to describe the extremal behavior of several common time series models with regular or irregular time sampling. The method is then used to analyze Hs data in the North Atlantic Ocean. The results indicate that it is possible to derive realistic estimates of the extremal properties of Hs from satellite data, despite its complex space--time sampling.

stat.AP

Gaussian linear state-space model for wind fields in the North-East Atlantic

A space-time model for wind fields is proposed. It aims at simulating realistic wind conditions with a focus on reproducing the space-time motions of the meteorological systems. A Gaussian linear state-space model is used where the latent state may be interpreted as regional wind condition and the observation equation links regional and local scales. Parameter estimation is performed by combining a method of moment and the EM algorithm whose performances are discussed using simulation studies. The model is fitted to 6-hourly reanalysis data in the North-East Atlantic. It is shown that the fitted model is interpretable and provide a good description of important properties of the space-time covariance function of the data, such as the non full-symmetry induced by prevailing flows in this area.

stat.ME

Modeling the Coastal Ocean over a Time Period of Several Weeks

From a scale analysis of hydrodynamic phenomena having a significant action on the drift of an object in coastal ocean waters, we deduce equations modeling the associated hydrodynamic fields over a time period of several weeks. These models are essentially non linear hyperbolic systems of PDE involving a small parameter. Then from the models we extract a simplified and nevertheless typical one for which we prove that its classical solution exists on a time interval which is independent of the small parameter. We then show that the solution weak-* converges as the small parameter goes to zero and we characterize the equation satisfied by the weak-* limit

math.AP

Long term object drift in the ocean with tide and wind

In this paper, we propose a new method to forecast the drift of objects in near coastal ocean on a period of several weeks. The proposed approach consists in estimating the probability of events linked to the drift using Monte Carlo simulations. It couples an averaging method which permits to decrease the computational cost and a statistical method in order to take into account the variability of meteorological loading factors.

math.NA