SearcharxivSearch

arXiv subjects

Olivier Roustant

Publications and source records attributed to Olivier Roustant.

At least 19 recordsLinked to original sources

A methodology for creating multidisciplinary design optimization benchmark problems from optimization ones

Benchmark problems with known solutions play a central role in the assessment of optimization algorithms. While mono-disciplinary optimization benefits from a rich collection of such problems, multidisciplinary design optimization (MDO) lacks equivalent resources: existing MDO benchmarks are scarce, rarely scalable, and their solutions are generally not known theoretically. In this paper, we propose a systematic methodology to transform any mono-disciplinary optimization problem with a known solution into a family of parametric MDO problems sharing that same solution. The construction relies on two key ingredients: a set of coupling equations that introduce interdependencies between disciplines, and a link function that eliminates the coupling variables and recovers the original mono-disciplinary problem. Theoretical conditions guaranteeing the equivalence between the two problems are established. The methodology is agnostic to the number of disciplines and variable dimensions, making it naturally suited for scalability studies. As an illustration, we construct a family of scalable MDO Rosenbrock problems and use them to benchmark two MDO coupling algorithms, namely the Jacobi and Gauss-Seidel schemes, across varying problem sizes. The proposed framework opens a systematic route to generating MDO benchmarks of arbitrary scale and complexity from the extensive catalog of existing mono-disciplinary test problems.

math.OC

Physics-informed, boundary-constrained Gaussian process regression for the reconstruction of fluid flow fields

Gaussian process regression techniques have been used in fluid mechanics for the reconstruction of flow fields from a reduction-of-dimension perspective. A main ingredient in this setting is the construction of adapted covariance functions, or kernels, to obtain such estimates. In this paper, we present a general method for constraining a prescribed Gaussian process on an arbitrary compact set. The kernel of the pre-defined process must be at least continuous and may include other information about the studied phenomenon. This general boundary-constraining framework can be implemented with high flexibility for a broad range of engineering applications. From this, we derive physics-informed kernels for simulating two-dimensional velocity fields of an incompressible (divergence-free) flow around aerodynamic profiles. These kernels allow to define Gaussian process priors satisfying the incompressibility condition and the prescribed boundary conditions along the profile in a continuous manner. We describe an adapted numerical method for the boundary-constraining procedure parameterized by a measure on the compact set. The relevance of the methodology and performances are illustrated by numerical simulations of flows around a cylinder and a NACA 0412 airfoil profile, for which no observation at the boundary is needed at all.

physics.flu-dyn

Multifidelity Gaussian process regression for solving nonlinear partial differential equations

Solving nonlinear partial differential equations (PDEs) using kernel methods offers a compelling alternative to traditional numerical solvers. However, the performance of these methods strongly depends on the choice of kernel. In this work, as the available information is inherently multifidelity, we propose a kernel learning approach based on cokriging, leveraging empirical information from multifidelity simulations. In the first step, we fit a differentiable non-stationary kernel to an empirical kernel obtained from low-fidelity simulations. In the second step, we derive a high-fidelity kernel with estimated hyperparameters, and construct a corresponding high-fidelity mean using the multifidelity framework. These components can then be used within a Gaussian process framework for solving PDEs. Finally, we demonstrate the performance of the proposed physics-informed method on the Burgers' equation.

stat.ML

Estimation and model errors in Gaussian-process-based Sensitivity Analysis of functional outputs

Global sensitivity analysis (GSA) of functional-output models is usually performed by combining statistical techniques, such as basis expansions, metamodeling and sampling based estimation of sensitivity indices. By neglecting truncation error from basis expansion, two main sources of errors propagate to the final sensitivity indices: the metamodeling related error and the sampling-based, or pick-freeze (PF), estimation error. This work provides an efficient algorithm to estimate these errors in the frame of Gaussian processes (GP), based on the approach of Le Gratiet et al. [16]. The proposed algorithm takes advantage of the fact that the number of basis coefficients of expanded model outputs is significantly smaller than output dimensions. Basis coefficients are fitted by GP models and multiple conditional GP trajectories are sampled. Then, vector-valued PF estimation is used to speed-up the estimation of Sobol indices and generalized sensitivity indices (GSI). We illustrate the methodology on an analytical test case and on an application in non-Newtonian hydraulics, modelling an idealized dam-break flow. Numerical tests show an improvement of 15 times in the computational time when compared to the application of Le Gratiet et al. [16] algorithm separately over each output dimension.

stat.ME

Simulation of extreme functionals in meteoceanic data: Application to surge evolution over tidal cycles

We investigate the influence of time-varying meteoceanic conditions on coastal flooding under the prism of rare events. Focusing on conditions observed over half tidal cycles, we observe that such data fall within the framework of functional extreme value theory, but violate standard assumptions due to temporal dependence and short-tailed behavior.a To address this, we propose a two-stage methodology. First, we introduce an autoregressive model to eliminate temporal dependence between cycles. Second, considering the model residuals, we adapt existing techniques based on Pareto processes. This allows us to build a simulator of extreme scenarios, by applying inverse transformations. These simulations depend on an initial time series, which can be randomly selected to tune the desired level of extremes. We validate the simulator performance by comparing simulated times series with observations, through several criteria, based on principal component analysis, extreme value analysis, and classification algorithms. The approach is applied to the surge data, on the G{â}vres site, located in southern Brittany, France.

stat.AP

Non-asymptotic confidence regions on RKHS. The Paley-Wiener and standard Sobolev space cases

We consider the problem of constructing a global, probabilistic, and non-asymptotic confidence region for an unknown function observed on a random design. The unknown function is assumed to lie in a reproducing kernel Hilbert space (RKHS). We show that this construction can be reduced to accurately estimating the RKHS norm of the unknown function. Our analysis primarily focuses both on the Paley-Wiener and on the standard Sobolev space settings.

math.ST

General reproducing properties in RKHS with application to derivative and integral operators

In this paper, we consider the reproducing property in Reproducing Kernel Hilbert Spaces (RKHS). We establish a reproducing property for the closure of the class of combinations of composition operators under minimal conditions. This allows to revisit the sufficient conditions for the reproducing property to hold for the derivative operator, as well as for the existence of the mean embedding function. These results provide a framework of application of the representer theorem for regularized learning algorithms that involve data for function values, gradients, or any other operator from the considered class.

math.ST

Fast pick-freeze estimation of Sobol' sensitivity maps using basis expansions

Global sensitivity analysis (GSA) aims at quantifying the contribution of input variables over the variability of model outputs. In the frame of functional outputs, a common goal is to compute sensitivity maps (SM), i.e sensitivity indices at each output dimension (e.g. time step for time series, or pixels for spatial outputs). In specific settings, some works have shown that the computation of Sobol' SM can be speeded up by using basis expansions employed for dimension reduction. However, how to efficiently compute such SM in a general setting has not received too much attention in the GSA literature.In this work, we propose fast computations of Sobol' SM using a general basis expansion, with a focus on statistical estimation. First, we write a closed-form expression of SM in function of the matrix-valued Sobol' index of the vector of basis coefficients. Secondly, we consider pick-freeze (PF) estimators, which have nice statistical properties (in terms of asymptotical efficiency) for Sobol' indices of any order. We provide similar basis-derived formulas for the PF estimator of Sobol' SM in function of the matrix-valued PF estimator of the vector of basis coefficients. We give the computational cost, and show that, compared to a dimension-wise approach, the computational gain is substantial and allows to calculate both SM and their associated bootstrap confidence bounds in a reasonable time. Finally, we illustrate the whole methodology on an analytical test case and on an application in non-Newtonian hydraulics, modelling an idealized dam-break flow.

math.ST

On one dimensional weighted Poincare inequalities for Global Sensitivity Analysis

One-dimensional Poincare inequalities are used in Global Sensitivity Analysis (GSA) to provide derivative-based upper bounds and approximations of Sobol indices. We add new perspectives by investigating weighted Poincare inequalities. Our contributions are twofold. In a first part, we provide new theoretical results for weighted Poincare inequalities, guided by GSA needs. We revisit the construction of weights from monotonic functions, providing a new proof from a spectral point of view. In this approach, given a monotonic function g, the weight is built such that g is the first non-trivial eigenfunction of a convenient diffusion operator. This allows us to reconsider the linear standard, i.e. the weight associated to a linear g. In particular, we construct weights that guarantee the existence of an orthonormal basis of eigenfunctions, leading to approximation of Sobol indices with Parseval formulas. In a second part, we develop specific methods for GSA. We study the equality case of the upper bound of a total Sobol index, and link the sharpness of the inequality to the proximity of the main effect to the eigenfunction. This leads us to theoretically investigate the construction of data-driven weights from estimators of the main effects when they are monotonic, another extension of the linear standard. Finally, we illustrate the benefits of using weights on a GSA study of two toy models and a real flooding application, involving the Poincare constant and/or the whole eigenbasis.

math.PR

Covariance models and Gaussian process regression for the wave equation. Application to related inverse problems

In this article, we consider the general task of performing Gaussian process regression (GPR) on pointwise observations of solutions of the 3 dimensional homogeneous free space wave equation.In a recent article, we obtained promising covariance expressions tailored to this equation: we now explore the potential applications of these formulas.We first study the particular cases of stationarity and radial symmetry, for which significant simplifications arise. We next show that the true-angle multilateration method for point source localization, as used in GPS systems, is naturally recovered by our GPR formulas in the limit of the small source radius. Additionally, we show that this GPR framework provides a new answer to the ill-posed inverse problem of reconstructing initial conditions for the wave equation from a limited number of sensors, and simultaneously enables the inference of physical parameters from these data. We finish by illustrating this ``physics informed'' GPR on a number of practical examples.

math.AP

Global sensitivity analysis using derivative-based sparse Poincaré chaos expansions

Variance-based global sensitivity analysis, in particular Sobol' analysis, is widely used for determining the importance of input variables to a computational model. Sobol' indices can be computed cheaply based on spectral methods like polynomial chaos expansions (PCE). Another choice are the recently developed Poincaré chaos expansions (PoinCE), whose orthonormal tensor-product basis is generated from the eigenfunctions of one-dimensional Poincaré differential operators. In this paper, we show that the Poincaré basis is the unique orthonormal basis with the property that partial derivatives of the basis form again an orthogonal basis with respect to the same measure as the original basis. This special property makes PoinCE ideally suited for incorporating derivative information into the surrogate modelling process. Assuming that partial derivative evaluations of the computational model are available, we compute spectral expansions in terms of Poincaré basis functions or basis partial derivatives, respectively, by sparse regression. We show on two numerical examples that the derivative-based expansions provide accurate estimates for Sobol' indices, even outperforming PCE in terms of bias and variance. In addition, we derive an analytical expression based on the PoinCE coefficients for a second popular sensitivity index, the derivative-based sensitivity measure (DGSM), and explore its performance as upper bound to the corresponding total Sobol' indices.

stat.CO

Characterization of the second order random fields subject to linear distributional PDE constraints

Let $L$ be a linear differential operator acting on functions defined over an open set $\mathcal{D}\subset \mathbb{R}^d$. In this article, we characterize the measurable second order random fields $U = (U(x))_{x\in\mathcal{D}}$ whose sample paths all verify the partial differential equation (PDE) $L(u) = 0$, solely in terms of their first two moments. When compared to previous similar results, the novelty lies in that the equality $L(u) = 0$ is understood in the sense of distributions, which is a powerful functional analysis framework mostly designed to study linear PDEs. This framework enables to reduce to the minimum the required differentiability assumptions over the first two moments of $(U(x))_{x\in\mathcal{D}}$ as well as over its sample paths in order to make sense of the PDE $L(U_ω)=0$. In view of Gaussian process regression (GPR) applications, we show that when $(U(x))_{x\in\mathcal{D}}$ is a Gaussian process (GP), the sample paths of $(U(x))_{x\in\mathcal{D}}$ conditioned on pointwise observations still verify the constraint $L(u)=0$ in the distributional sense. We finish by deriving a simple but instructive example, a GP model for the 3D linear wave equation, for which our theorem is applicable and where the previous results from the literature do not apply in general.

math.AP

Bayesian quadrature for $H^1(μ)$ with Poincaré inequality on a compact interval

Motivated by uncertainty quantification of complex systems, we aim at finding quadrature formulas of the form $\int_a^b f(x) dμ(x) = \sum_{i=1}^n w_i f(x_i)$ where $f$ belongs to $H^1(μ)$. Here, $μ$ belongs to a class of continuous probability distributions on $[a, b] \subset \mathbb{R}$ and $\sum_{i=1}^n w_i δ_{x_i}$ is a discrete probability distribution on $[a, b]$. We show that $H^1(μ)$ is a reproducing kernel Hilbert space with a continuous kernel $K$, which allows to reformulate the quadrature question as a Bayesian (or kernel) quadrature problem. Although $K$ has not an easy closed form in general, we establish a correspondence between its spectral decomposition and the one associated to Poincaré inequalities, whose common eigenfunctions form a $T$-system (Karlin and Studden, 1966). The quadrature problem can then be solved in the finite-dimensional proxy space spanned by the first eigenfunctions. The solution is given by a generalized Gaussian quadrature, which we call Poincaré quadrature. We derive several results for the Poincaré quadrature weights and the associated worst-case error. When $μ$ is the uniform distribution, the results are explicit: the Poincaré quadrature is equivalent to the midpoint (rectangle) quadrature rule. Its nodes coincide with the zeros of an eigenfunction and the worst-case error scales as $\frac{b-a}{2\sqrt{3}}n^{-1}$ for large $n$. By comparison with known results for $H^1(0,1)$, this shows that the Poincaré quadrature is asymptotically optimal. For a general $μ$, we provide an efficient numerical procedure, based on finite elements and linear programming. Numerical experiments provide useful insights: nodes are nearly evenly spaced, weights are close to the probability density at nodes, and the worst-case error is approximately $O(n^{-1})$ for large $n$.

math.ST

High-dimensional additive Gaussian processes under monotonicity constraints

We introduce an additive Gaussian process framework accounting for monotonicity constraints and scalable to high dimensions. Our contributions are threefold. First, we show that our framework enables to satisfy the constraints everywhere in the input space. We also show that more general componentwise linear inequality constraints can be handled similarly, such as componentwise convexity. Second, we propose the additive MaxMod algorithm for sequential dimension reduction. By sequentially maximizing a squared-norm criterion, MaxMod identifies the active input dimensions and refines the most important ones. This criterion can be computed explicitly at a linear cost. Finally, we provide open-source codes for our full framework. We demonstrate the performance and scalability of the methodology in several synthetic examples with hundreds of dimensions under monotonicity constraints as well as on a real-world flood application.

stat.ML

A comparison of mixed-variables Bayesian optimization approaches

Most real optimization problems are defined over a mixed search space where the variables are both discrete and continuous. In engineering applications, the objective function is typically calculated with a numerically costly black-box simulation.General mixed and costly optimization problems are therefore of a great practical interest, yet their resolution remains in a large part an open scientific question. In this article, costly mixed problems are approached through Gaussian processes where the discrete variables are relaxed into continuous latent variables. The continuous space is more easily harvested by classical Bayesian optimization techniques than a mixed space would. Discrete variables are recovered either subsequently to the continuous optimization, or simultaneously with an additional continuous-discrete compatibility constraint that is handled with augmented Lagrangians. Several possible implementations of such Bayesian mixed optimizers are compared. In particular, the reformulation of the problem with continuous latent variables is put in competition with searches working directly in the mixed space. Among the algorithms involving latent variables and an augmented Lagrangian, a particular attention is devoted to the Lagrange multipliers for which a local and a global estimation techniques are studied. The comparisons are based on the repeated optimization of three analytical functions and a beam design problem.

math.OC

Sequential construction and dimension reduction of Gaussian processes under constraints

Accounting for inequality constraints, such as boundedness, monotonicity or convexity, is challenging when modeling costly-to-evaluate black box functions. In this regard, finite-dimensional Gaussian process (GP) regression models bring a valuable solution, as they guarantee that the inequality constraints are satisfied everywhere. Nevertheless, these models are currently restricted to small dimensional situations (up to dimension 5). Addressing this issue, we introduce the MaxMod algorithm that sequentially inserts one-dimensional knots or adds active variables, thereby performing at the same time dimension reduction and efficient knot allocation. We prove the convergence of this algorithm. In intermediary steps of the proof, we propose the notion of multi-affine extension and study its properties. We also prove the convergence of finite-dimensional GPs, when the knots are not dense in the input space, extending the recent literature. With simulated and real data, we demonstrate that the MaxMod algorithm remains efficient in higher dimension (at least in dimension 20), and needs fewer knots than other constrained GP models from the state-of-the-art, to reach a given approximation error.

math.ST

Stochastic Processes Under Linear Differential Constraints : Application to Gaussian Process Regression for the 3 Dimensional Free Space Wave Equation

Let $P$ be a linear differential operator over $\mathcal{D} \subset \mathbb{R}^d$ and $U = (U_x)_{x \in \mathcal{D}}$ a second order stochastic process. In the first part of this article, we prove a new necessary and sufficient condition for all the trajectories of $U$ to verify the partial differential equation (PDE) $T(U) = 0$. This condition is formulated in terms of the covariance kernel of $U$. When compared to previous similar results, the novelty lies in that the equality $T(U) = 0$ is understood in the \textit{sense of distributions}, which is a relevant framework for PDEs. This theorem provides precious insights during the second part of this article, devoted to performing "physically informed" machine learning for the homogeneous 3 dimensional free space wave equation. We perform Gaussian process regression (GPR) on pointwise observations of a solution of this PDE. To do so, we propagate Gaussian processes (GP) priors over its initial conditions through the wave equation. We obtain explicit formulas for the covariance kernel of the propagated GP, which can then be used for GPR. We then explore the particular cases of radial symmetry and point source. For the former, we derive convolution-free GPR formulas; for the latter, we show a direct link between GPR and the classical triangulation method for point source localization used in GPS systems. Additionally, this Bayesian framework provides a new answer for the ill-posed inverse problem of reconstructing initial conditions for the wave equation with a limited number of sensors, and simultaneously enables the inference of physical parameters from these data. Finally, we illustrate this physically informed GPR on a number of practical examples.

math.ST

Approximating Gaussian Process Emulators with Linear Inequality Constraints and Noisy Observations via MC and MCMC

Adding inequality constraints (e.g. boundedness, monotonicity, convexity) into Gaussian processes (GPs) can lead to more realistic stochastic emulators. Due to the truncated Gaussianity of the posterior, its distribution has to be approximated. In this work, we consider Monte Carlo (MC) and Markov Chain Monte Carlo (MCMC) methods. However, strictly interpolating the observations may entail expensive computations due to highly restrictive sample spaces. Furthermore, having (constrained) GP emulators when data are actually noisy is also of interest for real-world implementations. Hence, we introduce a noise term for the relaxation of the interpolation conditions, and we develop the corresponding approximation of GP emulators under linear inequality constraints. We show with various toy examples that the performance of MC and MCMC samplers improves when considering noisy observations. Finally, on 2D and 5D coastal flooding applications, we show that more flexible and realistic GP implementations can be obtained by considering noise effects and by enforcing the (linear) inequality constraints.

stat.ML