Searcharxiv⌕ Search

arXiv subjects

Liang Yan

Publications and source records attributed to Liang Yan.

27 records · Page 2Linked to original sources

A Less Uncertain Sampling-Based Method of Batch Bayesian Optimization

This paper presents a method called sampling-computation-optimization (SCO) to design batch Bayesian optimization. SCO does not construct new high-dimensional acquisition functions but samples from the existing one-site acquisition function to obtain several candidate samples. To reduce the uncertainty of the sampling, the general discrepancy is computed to compare these samples. Finally, the genetic algorithm and switch algorithm are used to optimize the design. Several strategies are used to reduce the computational burden in the SCO. From the numerical results, the SCO designs were less uncertain than those of other sampling-based methods. As for application in batch Bayesian optimization, SCO can find a better solution when compared with other batch methods in the same dimension and batch size. In addition, it is also flexible and can be adapted to different one-site methods. Finally, a complex experimental case is given to illustrate the application value and scenario of SCO method.

math.OC↗

An acceleration strategy for randomize-then-optimize sampling via deep neural networks

Randomize-then-optimize (RTO) is widely used for sampling from posterior distributions in Bayesian inverse problems. However, RTO may be computationally intensive for complexity problems due to repetitive evaluations of the expensive forward model and its gradient. In this work, we present a novel strategy to substantially reduce the computation burden of RTO by using a goal-oriented deep neural networks (DNN) surrogate approach. In particular, the training points for the DNN-surrogate are drawn from a local approximated posterior distribution, and it is shown that the resulting algorithm can provide a flexible and efficient sampling algorithm, which converges to the direct RTO approach. We present a Bayesian inverse problem governed by a benchmark elliptic PDE to demonstrate the computational accuracy and efficiency of our new algorithm (i.e., DNN-RTO). It is shown that with our algorithm, one can significantly outperform the traditional RTO.

math.NA↗

Optimal design for kernel interpolation: applications to uncertainty quantification

The paper is concerned with classic kernel interpolation methods, in addition to approximation methods that are augmented by gradient measurements. To apply kernel interpolation using radial basis functions (RBFs) in a stable way, we propose a type of quasi-optimal interpolation points, searching from a large set of \textit{candidate} points, using a procedure similar to designing Fekete points or power function maximizing points that use pivot from a Cholesky decomposition. The proposed quasi-optimal points result in a smaller condition number, and thus mitigates the instability of the interpolation procedure when the number of points becomes large. Applications to parametric uncertainty quantification are presented, and it is shown that the proposed interpolation method can outperform sparse grid methods in many interesting cases. We also demonstrate the new procedure can be applied to constructing gradient-enhanced Gaussian process emulators.

math.NA↗

Stein variational gradient descent with local approximations

Bayesian computation plays an important role in modern machine learning and statistics to reason about uncertainty. A key computational challenge in Bayesian inference is to develop efficient techniques to approximate, or draw samples from posterior distributions. Stein variational gradient decent (SVGD) has been shown to be a powerful approximate inference algorithm for this issue. However, the vanilla SVGD requires calculating the gradient of the target density and cannot be applied when the gradient is unavailable or too expensive to evaluate. In this paper we explore one way to address this challenge by the construction of a local surrogate for the target distribution in which the gradient can be obtained in a much more computationally feasible manner. More specifically, we approximate the forward model using a deep neural network (DNN) which is trained on a carefully chosen training set, which also determines the quality of the surrogate. To this end, we propose a general adaptation procedure to refine the local approximation online without destroying the convergence of the resulting SVGD. This significantly reduces the computational cost of SVGD and leads to a suite of algorithms that are straightforward to implement. The new algorithm is illustrated on a set of challenging Bayesian inverse problems, and numerical experiments demonstrate a clear improvement in performance and applicability of standard SVGD.

math.NA↗

An adaptive surrogate modeling based on deep neural networks for large-scale Bayesian inverse problems

In Bayesian inverse problems, surrogate models are often constructed to speed up the computational procedure, as the parameter-to-data map can be very expensive to evaluate. However, due to the curse of dimensionality and the nonlinear concentration of the posterior, traditional surrogate approaches (such us the polynomial-based surrogates) are still not feasible for large scale problems. To this end, we present in this work an adaptive multi-fidelity surrogate modeling framework based on deep neural networks (DNNs), motivated by the facts that the DNNs can potentially handle functions with limited regularity and are powerful tools for high dimensional approximations. More precisely, we first construct offline a DNNs-based surrogate according to the prior distribution, and then, this prior-based DNN-surrogate will be adaptively \& locally refined online using only a few high-fidelity simulations. In particular, in the refine procedure, we construct a new shallow neural network that view the previous constructed surrogate as an input variable -- yielding a composite multi-fidelity neural network approach. This makes the online computational procedure rather efficient. Numerical examples are presented to confirm that the proposed approach can obtain accurate posterior information with a limited number of forward simulations.

math.NA↗

A non-intrusive reduced basis EKI for time-fractional diffusion inverse problems

In this study, we consider an ensemble Kalman inversion (EKI) for the numerical solution of time-fractional diffusion inverse problems (TFDIPs). Computational challenges in the EKI arise from the need for repeated evaluations of the forward model. We address this challenge by introducing a non-intrusive reduced basis (RB) method for constructing surrogate models to reduce computational cost. In this method, a reduced basis is extracted from a set of full-order snapshots by the proper orthogonal decomposition (POD), and a doubly stochastic radial basis function (DSRBF) is used to learn the projection coefficients. The DSRBF is carried out in the offline stage with a stochastic leave-one-out cross-validation algorithm to select the shape parameter, and the outputs for new parameter values can be obtained rapidly during the online stage. Due to the complete decoupling of the offline and online stages, the proposed non-intrusive RB method -- referred to as POD-DSRBF -- provides a powerful tool to accelerate the EKI approach for TFDIPs. We demonstrate the practical performance of the proposed strategies through two nonlinear time-fractional diffusion inverse problems. The numerical results indicate that the new algorithm can achieve significant computational gains without sacrificing accuracy.

math.NA↗

Adaptive multi-fidelity polynomial chaos approach to Bayesian inference in inverse problems

The polynomial chaos (PC) expansion has been widely used as a surrogate model in the Bayesian inference to speed up the Markov chain Monte Carlo (MCMC) calculations. However, the use of a PC surrogate introduces the modeling error, that may severely distort the estimate of the posterior distribution. This error can be corrected by increasing the order of the PC expansion, but the cost for building the surrogate may increase dramatically. In this work, we seek to address this challenge by proposing an adaptive procedure to construct a multi-fidelity PC surrogate. This new strategy combines (a large number of) low-fidelity surrogate model evaluations and (a small number of) high-fidelity model evaluations, yielding an adaptive multi-fidelity approach. Here the low-fidelity surrogate is chosen as the prior-based PC surrogate, while the high-fidelity model refers to the true forward model. The key idea is to construct and refine the multi-fidelity approach over a sequence of samples adaptively determined form data, so that the approximation can eventually concentrate to the posterior distribution. We illustrate the performance of the proposed strategy through two nonlinear inverse problems. It is shown that the proposed adaptive multi-fidelity approach can improve significantly the accuracy, yet without a dramatic increase in the computational complexity. The numerical results also indicate that our new algorithm can enhance the efficiency by several orders of magnitude compared to a standard MCMC approach using only the true forward model.

math.NA↗

An adaptive multi-fidelity PC-based ensemble Kalman inversion for inverse problems

The ensemble Kalman inversion (EKI), as a derivative-free methodology, has been widely used in the parameter estimation of inverse problems. Unfortunately, its cost may become moderately large for systems described by high dimensional nonlinear PDEs, as EKI requires a relatively large ensemble size to guarantee its performance. In this paper, we propose an adaptive multi-fidelity polynomial chaos (PC) based EKI technique to address this challenge. Our new strategy combines a large number of low-order PC surrogate model evaluations and a small number of high-fidelity forward model evaluations, yielding a multi-fidelity approach. Especially, we present a new approach that adaptively constructs and refines a multi-fidelity PC surrogate during the EKI simulation. Since the forward model evaluations are only required for updating the low-order multi-fidelity PC model, whose number can be much smaller than the total ensemble size of the classic EKI, the entire computational costs are thus significantly reduced. The new algorithm was tested through the two-dimensional time fractional inverse diffusion problems and demonstrated great effectiveness in comparison with PC based EKI and classic EKI.

math.NA↗

Weighted approximate Fekete points: Sampling for least-squares polynomial approximation

We propose and analyze a weighted greedy scheme for computing deterministic sample configurations in multidimensional space for performing least-squares polynomial approximations on $L^2$ spaces weighted by a probability density function. Our procedure is a particular weighted version of the approximate Fekete points method, with the weight function chosen as the (inverse) Christoffel function. Our procedure has theoretical advantages: when linear systems with optimal condition number exist, the procedure finds them. In the one-dimensional setting with any density function, our greedy procedure almost always generates optimally-conditioned linear systems. Our method also has practical advantages: our procedure is impartial to compactness of the domain of approximation, and uses only pivoted linear algebraic routines. We show through numerous examples that our sampling design outperforms competing randomized and deterministic designs when the domain is both low and high dimensional.

math.NA↗