SearcharxivSearch

arXiv subjects

Yann Richet

Publications and source records attributed to Yann Richet.

5 recordsLinked to original sources

Augmented Quantization: Mixture Models for Risk-Oriented Sensitivity Analysis

A central question in risk analysis is to identify the factors that drive the system toward a specific hazardous outcome, such as the exceedance of a given threshold. When relying on numerical simulators, we propose to study the distribution of the inputs, transformed into uniform variables via their cumulative distributions, conditionally on the occurrence of the hazardous event. To represent this multivariate conditional distribution for sensitivity analysis, we introduce an original quantization approach based on estimating a mixture of Dirac and local uniform distributions. For each marginal of this mixture, a Dirac component indicates a strong influence of the corresponding variable, whereas a uniform component with wide support reflects weak influence. A notable advantage of this method is its ability to identify the regions of the input space that most strongly influence the occurrence of the risk event, while also capturing the joint effects of multiple variables. However, learning mixture models typically relies on likelihood-based methods, which are not well suited to mixtures involving singular or Dirac components. To address this, we propose an \emph{Augmented Quantization} method, a reformulation of the classical quantization problem based on the p-Wasserstein distance, which can be computed in very general distribution spaces. The performance of Augmented Quantization in estimating such mixture models is first demonstrated on analytical toy problems, and then applied to sensitivity analysis of both an analytical function and a practical flooding case study on a section of the Loire River.

stat.AP

FunQuant: A R package to perform quantization in the context of rare events and time-consuming simulations

Quantization summarizes continuous distributions by calculating a discrete approximation. Among the widely adopted methods for data quantization is Lloyd's algorithm, which partitions the space into Vorono\"i cells, that can be seen as clusters, and constructs a discrete distribution based on their centroids and probabilistic masses. Lloyd's algorithm estimates the optimal centroids in a minimal expected distance sense, but this approach poses significant challenges in scenarios where data evaluation is costly, and relates to rare events. Then, the single cluster associated to no event takes the majority of the probability mass. In this context, a metamodel is required and adapted sampling methods are necessary to increase the precision of the computations on the rare clusters.

stat.CO

Quantizing rare random maps: application to flooding visualization

Visualization is an essential operation when assessing the risk of rare events such as coastal or river floodings. The goal is to display a few prototype events that best represent the probability law of the observed phenomenon, a task known as quantization. It becomes a challenge when data is expensive to generate and critical events are scarce, like extreme natural hazard. In the case of floodings, each event relies on an expensive-to-evaluate hydraulic simulator which takes as inputs offshore meteo-oceanic conditions and dyke breach parameters to compute the water level map. In this article, Lloyd's algorithm, which classically serves to quantize data, is adapted to the context of rare and costly-to-observe events. Low probability is treated through importance sampling, while Functional Principal Component Analysis combined with a Gaussian process deal with the costly hydraulic simulations. The calculated prototype maps represent the probability distribution of the flooding events in a minimal expected distance sense, and each is associated to a probability mass. The method is first validated using a 2D analytical model and then applied to a real coastal flooding scenario. The two sources of error, the metamodel and the importance sampling, are evaluated to quantify the precision of the method.

stat.AP

Sequential Design of Mixture Experiments with an Empirically Determined Input Domain and an Application to Burn-up Credit Penalization of Nuclear Fuel Rods

This paper proposes a sequential design for maximizing a stochastic computer simulator output, y(x), over an unknown optimization domain. The training data used to estimate the optimization domain are a set of (historical) inputs, often from a physical system modeled by the simulator. Two methods are provided for estimating the simulator input domain. An extension of the well-known efficient global optimization algorithm is presented to maximize y(x). The domain estimation/maximization procedure is applied to two readily understood analytic examples. It is also used to solve a problem in nuclear safety by maximizing the k-effective "criticality coefficient" of spent fuel rods, considered as one-dimensional heterogeneous fissile media. One of the two domain estimation methods relies on expertise-type constraints. We show that these constraints, initially chosen to address the spent fuel rod example, are robust in that they also lead to good results in the second analytic optimization example. Of course, in other applications, it could be necessary to design alternative constraints that are more suitable for these applications.

stat.ME

Adaptive Design of Experiments for Conservative Estimation of Excursion Sets

We consider the problem of estimating the set of all inputs that leads a system to some particular behavior. The system is modeled by an expensive-to-evaluate function, such as a computer experiment, and we are interested in its excursion set, i.e. the set of points where the function takes values above or below some prescribed threshold. The objective function is emulated with a Gaussian Process (GP) model based on an initial design of experiments enriched with evaluation results at (batch-)sequentially determined input points. The GP model provides conservative estimates for the excursion set, which control false positives while minimizing false negatives. We introduce adaptive strategies that sequentially select new evaluations of the function by reducing the uncertainty on conservative estimates. Following the Stepwise Uncertainty Reduction approach we obtain new evaluations by minimizing adapted criteria. Tractable formulae for the conservative criteria are derived, which allow more convenient optimization. The method is benchmarked on random functions generated under the model assumptions in different scenarios of noise and batch size. We then apply it to a reliability engineering test case. Overall, the proposed strategy of minimizing false negatives in conservative estimation achieves competitive performance both in terms of model-based and model-free indicators.

stat.ME