SearcharxivSearch

arXiv subjects

Guillaume Damblin

Publications and source records attributed to Guillaume Damblin.

9 recordsLinked to original sources

A Gaussian process and linear-based framework for computing cut distributions in modular Bayesian calibration of two chained computer models

Computer models are widely used in science and engineering to simulate complex systems. However, these models are affected by several sources of uncertainty, which may limit their use for decision making in risk management. We present a Bayesian approach for quantifying parameter uncertainty in a chain of two computer models motivated by multiphysics simulations in the nuclear field. Part of the inputs of a downstream model parametrized by $θ\in \mathbb{R}^p$ come from the outputs of an upstream model parametrized by $λ\in \mathbb{R}^q$. Usually, the joint posterior distribution of $(θ, λ)$ would be obtained by applying Bayes' theorem using the experimental observations of both models. However, when the observations of the downstream model are too indirect to provide informative inference on $λ$, it may be preferable to compute a modular posterior distribution of $(θ, λ)$, referred to as the \emph{cut distribution}. Assuming that the posterior distribution of $λ$ has been previously estimated from observations of the upstream model only, we aim to compute the posterior distribution of $θ$ conditional on $λ$ using observations from the downstream model. To this end, we propose a Gaussian-process and linear-based framework to estimate the functional dependence between $θ$ and $λ$, denoted by $θ(λ)$, where each component is modeled as a realization of a Gaussian process. As the downstream model is approximated by a linear function of $θ(λ)$, Bayesian conjugacy allows us to derive a Gaussian posterior predictive distribution of $θ(λ)$ for any realization of $λ$. The effectiveness of the method is illustrated through several synthetic examples, and we highlight how variations in $λ$ impact the predictive distribution of the chained simulation.

stat.CO

Multivariate Bayesian Last Layer for Regression with Uncertainty Quantification and Decomposition

We present new Bayesian Last Layer neural network models in the setting of multivariate regression under heteroscedastic noise, and propose EM algorithms for parameter learning. Bayesian modeling of a neural network's final layer has the attractive property of uncertainty quantification with a single forward pass. The proposed framework is capable of disentangling the aleatoric and epistemic uncertainty, and can be used to enhance a canonically trained deep neural network with uncertainty-aware capabilities.

stat.ML

Revisiting Tensor Basis Neural Networks for Reynolds stress modeling: application to plane channel and square duct flows

Several Tensor Basis Neural Network (TBNN) frameworks aimed at enhancing turbulence RANS modeling have recently been proposed in the literature as data-driven constitutive models for systems with known invariance properties. However, persistent ambiguities remain regarding the physical adequacy of applying the General Eddy Viscosity Model (GEVM). This work aims at investigating this aspect in an a priori stage for better predictions of the Reynolds stress anisotropy tensor, while preserving the Galilean and rotational invariances. In particular, we propose a general framework providing optimal tensor basis models for two types of canonical flows: Plane Channel Flow (PCF) and Square Duct Flow (SDF). Subsequently, deep neural networks based on these optimal models are trained using state-of-the-art strategies to achieve a balanced and physically sound prediction of the full anisotropy tensor. A priori results obtained by the proposed framework are in very good agreement with the reference DNS data. Notably, our shallow network with three layers provides accurate predictions of the anisotropy tensor for PCF at unobserved friction Reynolds numbers, both in interpolation and extrapolation scenarios. Learning the SDF case is more challenging because of its physical nature and a lack of training data at various regimes. We propose to alleviate this problem based on Transfer Learning (TL). To more efficiently generalize to an unseen intermediate $\mathrm{Re}_τ$ regime, we take advantage of our prior knowledge acquired from a training with a larger and wider dataset. Our results indicate the potential of the developed network model, and demonstrate the feasibility and efficiency of the TL process in terms of training data size and training time. Based on these results, we believe there is a promising future by integrating these neural networks into an adapted in-house RANS solver.

physics.flu-dyn

A generalization of the CIRCE method for quantifying input model uncertainty in presence of several groups of experiments

The semi-empirical nature of best-estimate models closing the balance equations of thermal-hydraulic (TH) system codes is well-known as a significant source of uncertainty for accuracy of output predictions. This uncertainty, called model uncertainty, is usually represented by multiplicative (log-)Gaussian variables whose estimation requires solving an inverse problem based on a set of adequately chosen real experiments. One method from the TH field, called CIRCE, addresses it. We present in the paper a generalization of this method to several groups of experiments each having their own properties, including different ranges for input conditions and different geometries. An individual (log-)Gaussian distribution is therefore estimated for each group in order to investigate whether the model uncertainty is homogeneous between the groups, or should depend on the group. To this end, a multi-group CIRCE is proposed where a variance parameter is estimated for each group jointly to a mean parameter common to all the groups to preserve the uniqueness of the best-estimate model. The ECME algorithm for Maximum Likelihood Estimation is adapted to the latter context, then applied to relevant demonstration cases. Finally, it is tested on a practical case to assess the uncertainty of critical mass flow assuming two groups due to the difference of geometry between the experimental setups.

stat.ME

Reynolds Stress Anisotropy Tensor Predictions for Turbulent Channel Flow using Neural Networks

The Reynolds-Averaged Navier-Stokes (RANS) approach remains a backbone for turbulence modeling due to its high cost-effectiveness. Its accuracy is largely based on a reliable Reynolds stress anisotropy tensor closure model. There has been an amount of work aiming at improving traditional closure models, while they are still not satisfactory to some complex flow configurations. In recent years, advances in computing power have opened up a new way to address this problem: the machine-learning-assisted turbulence modeling. In this paper, we employ neural networks to fully predict the Reynolds stress anisotropy tensor of turbulent channel flows at different friction Reynolds numbers, for both interpolation and extrapolation scenarios. Several generic neural networks of Multi-Layer Perceptron (MLP) type are trained with different input feature combinations to acquire a complete grasp of the role of each parameter. The best performance is yielded by the model with the dimensionless mean streamwise velocity gradient $α$, the dimensionless wall distance $y^+$ and the friction Reynolds number $\mathrm{Re}_τ$ as inputs. A deeper theoretical insight into the Tensor Basis Neural Network (TBNN) clarifies some remaining ambiguities found in the literature concerning its application of Pope's general eddy viscosity model. We emphasize the sensitivity of the TBNN on the constant tensor $\textbf{T}^{*(0)}$ upon the turbulent channel flow data set, and newly propose a generalized $\textbf{T}^{*(0)}$, which considerably enhances its performance. Through comparison between the MLP and the augmented TBNN model with both $\{α, y^+, \mathrm{Re}_τ\}$ as input set, it is concluded that the former outperforms the latter and provides excellent interpolation and extrapolation predictions of the Reynolds stress anisotropy tensor in the specific case of turbulent channel flow.

physics.flu-dyn

Adaptive use of replicated Latin Hypercube Designs for computing Sobol' sensitivity indices

As recently pointed out in the field of Global Sensitivity Analysis (GSA) of computer simulations, the use of replicated Latin Hypercube Designs (rLHDs) is a cost-saving alternative to regular Monte Carlo sampling to estimate first-order Sobol' indices. Indeed, two rLHDs are sufficient to compute the whole set of those indices regardless of the number of input variables. This relies on a permutation trick which, however, only works within the class of estimators called Oracle 2. In the present paper, we show that rLHDs are still beneficial to another class of estimators, called Oracle 1, which often outperforms Oracle 2 for estimating small and moderate indices. Even though unlike Oracle 2 the computation cost of Oracle 1 depends on the input dimension, the permutation trick can be applied to construct an averaged (triple) Oracle 1 estimator whose great accuracy is presented on a numerical example. Thus, we promote an adaptive rLHDs-based Sobol' sensitivity analysis where the first stage is to compute the whole set of first-order indices by Oracle 2. If needed, the accuracy of small and moderate indices can then be reevaluated by the averaged Oracle 1 estimators. This strategy, cost-saving and guaranteeing the accuracy of estimates, is applied to a computer model from the nuclear field.

stat.CO

Bayesian inference and non-linear extensions of the CIRCE method for quantifying the uncertainty of closure relationships integrated into thermal-hydraulic system codes

Uncertainty Quantification of closure relationships integrated into thermal-hydraulic system codes is a critical prerequisite in applying the Best-Estimate Plus Uncertainty (BEPU) methodology for nuclear safety and licensing processes.The purpose of the CIRCE method is to estimate the (log)-Gaussian probability distribution of a multiplicative factor applied to a reference closure relationship in order to assess its uncertainty. Even though this method has been implemented with success in numerous physical scenarios, it can still suffer from substantial limitations such as the linearity assumption and the difficulty of properly taking into account the inherent statistical uncertainty. In the paper, we will extend the CIRCE method in two aspects. On the one hand, we adopt the Bayesian setting putting prior probability distributions on the parameters of the (log)-Gaussian distribution. The posterior distribution of the parameters is then computed with respect to an experimental database by means of Markov Chain Monte Carlo (MCMC) algorithms. On the other hand, we tackle the more general setting where the simulations do not move linearly against the multiplicative factor(s). MCMC algorithms then become time-prohibitive when the thermal-hydraulic simulations exceed a few minutes. This handicap is overcome by using Gaussian process (GP) emulators which can yield both reliable and fast predictions of the simulations. The GP-based MCMC algorithms will be applied to quantify the uncertainty of two condensation closure relationships at a safety injection with respect to a database of experimental tests. The thermal-hydraulic simulations will be run with the CATHARE 2 computer code.

stat.CO

Adaptive numerical designs for the calibration of computer codes

Making good predictions of a physical system using a computer code requires the inputs to be carefully specified. Some of these inputs called control variables have to reproduce physical conditions whereas other inputs, called parameters, are specific to the computer code and most often uncertain. The goal of statistical calibration consists in estimating these parameters with the help of a statistical model which links the code outputs with the field measurements. In a Bayesian setting, the posterior distribution of these parameters is normally sampled using MCMC methods. However, they are impractical when the code runs are high time-consuming. A way to circumvent this issue consists of replacing the computer code with a Gaussian process emulator, then sampling a cheap-to-evaluate posterior distribution based on it. Doing so, calibration is subject to an error which strongly depends on the numerical design of experiments used to fit the emulator. We aim at reducing this error by building a proper sequential design by means of the Expected Improvement criterion. Numerical illustrations in several dimensions assess the efficiency of such sequential strategies.

stat.CO

Numerical studies of space filling designs: optimization of Latin Hypercube Samples and subprojection properties

Quantitative assessment of the uncertainties tainting the results of computer simulations is nowadays a major topic of interest in both industrial and scientific communities. One of the key issues in such studies is to get information about the output when the numerical simulations are expensive to run. This paper considers the problem of exploring the whole space of variations of the computer model input variables in the context of a large dimensional exploration space. Various properties of space filling designs are justified: interpoint-distance, discrepancy, minimum spanning tree criteria. A specific class of design, the optimized Latin Hypercube Sample, is considered. Several optimization algorithms, coming from the literature, are studied in terms of convergence speed, robustness to subprojection and space filling properties of the resulting design. Some recommendations for building such designs are given. Finally, another contribution of this paper is the deep analysis of the space filling properties of the design 2D-subprojections.

math.ST