Searcharxiv⌕ Search

arXiv subjects

Guang Lin

Publications and source records attributed to Guang Lin.

At least 181 records · Page 10Linked to original sources

Trimmed Ensemble Kalman Filter for Nonlinear and Non-Gaussian Data Assimilation Problems

We study the ensemble Kalman filter (EnKF) algorithm for sequential data assimilation in a general situation, that is, for nonlinear forecast and measurement models with non-additive and non-Gaussian noises. Such applications traditionally force us to choose between inaccurate Gaussian assumptions that permit efficient algorithms (e.g., EnKF), or more accurate direct sampling methods which scale poorly with dimension (e.g., particle filters, or PF). We introduce a trimmed ensemble Kalman filter (TEnKF) which can interpolate between the limiting distributions of the EnKF and PF to facilitate adaptive control over both accuracy and efficiency. This is achieved by introducing a trimming function that removes non-Gaussian outliers that introduce errors in the correlation between the model and observed forecast, which otherwise prevent the EnKF from proposing accurate forecast updates. We show for specific trimming functions that the TEnKF exactly reproduces the limiting distributions of the EnKF and PF. We also develop an adaptive implementation which provides control of the effective sample size and allows the filter to overcome periods of increased model nonlinearity. This algorithm allow us to demonstrate substantial improvements over the traditional EnKF in convergence and robustness for the nonlinear Lorenz-63 and Lorenz-96 models.

stat.ME↗

Inverse modeling of hydrologic systems with adaptive multi-fidelity Markov chain Monte Carlo simulations

Markov chain Monte Carlo (MCMC) simulation methods are widely used to assess parametric uncertainties of hydrologic models conditioned on measurements of observable state variables. However, when the model is CPU-intensive and high-dimensional, the computational cost of MCMC simulation will be prohibitive. In this situation, a CPU-efficient while less accurate low-fidelity model (e.g., a numerical model with a coarser discretization, or a data-driven surrogate) is usually adopted. Nowadays, multi-fidelity simulation methods that can take advantage of both the efficiency of the low-fidelity model and the accuracy of the high-fidelity model are gaining popularity. In the MCMC simulation, as the posterior distribution of the unknown model parameters is the region of interest, it is wise to distribute most of the computational budget (i.e., the high-fidelity model evaluations) therein. Based on this idea, in this paper we propose an adaptive multi-fidelity MCMC algorithm for efficient inverse modeling of hydrologic systems. In this method, we evaluate the high-fidelity model mainly in the posterior region through iteratively running MCMC based on a Gaussian process (GP) system that is adaptively constructed with multi-fidelity simulation. The error of the GP system is rigorously considered in the MCMC simulation and gradually reduced to a negligible level in the posterior region. Thus, the proposed method can obtain an accurate estimate of the posterior distribution with a small number of the high-fidelity model evaluations. The performance of the proposed method is demonstrated by three numerical case studies in inverse modeling of hydrologic systems.

math.OC↗

An iterative local updating ensemble smoother for estimation and uncertainty assessment of hydrologic model parameters with multimodal distributions

Ensemble smoother (ES) has been widely used in inverse modeling of hydrologic systems. However, for problems where the distribution of model parameters is multimodal, using ES directly would be problematic. One popular solution is to use a clustering algorithm to identify each mode and update the clusters with ES separately. However, this strategy may not be very efficient when the dimension of parameter space is high or the number of modes is large. Alternatively, we propose in this paper a very simple and efficient algorithm, i.e., the iterative local updating ensemble smoother (ILUES), to explore multimodal distributions of model parameters in nonlinear hydrologic systems. The ILUES algorithm works by updating local ensembles of each sample with ES to explore possible multimodal distributions. To achieve satisfactory data matches in nonlinear problems, we adopt an iterative form of ES to assimilate the measurements multiple times. Numerical cases involving nonlinearity and multimodality are tested to illustrate the performance of the proposed method. It is shown that overall the ILUES algorithm can well quantify the parametric uncertainties of complex hydrologic models, no matter whether the multimodal distribution exists.

math.OC↗

Turbulence Generation from a stochastic wavelet model

This research presents a new turbulence generation method based on stochastic wavelets and tests its various properties in both homogeneous and inhomogeneous turbulence. Turbulence field can be generated with less basis compared to previous synthetic Fourier methods. Adaptive generation of inhomogeneous turbulence is achieved by scale reduction algorithm and lead to smaller computation cost. The generated turbulence shows good agreement with input data and theoretical results.

math.NA↗

On the Bayesian calibration of expensive computer models with input dependent parameters

Computer models, aiming at simulating a complex real system, are often calibrated in the light of data to improve performance. Standard calibration methods assume that the optimal values of calibration parameters are invariant to the model inputs. In several real world applications where models involve complex parametrizations whose optimal values depend on the model inputs, such an assumption can be too restrictive and may lead to misleading results. We propose a fully Bayesian methodology that produces input dependent optimal values for the calibration parameters, as well as it characterizes the associated uncertainties via posterior distributions. Central to methodology is the idea of formulating the calibration parameter as a step function whose uncertain structure is modeled properly via a binary treed process. Our method is particularly suitable to address problems where the computer model requires the selection of a sub-model from a set of competing ones, but the choice of the `best' sub-model may change with the input values. The method produces a selection probability for each sub-model given the input. We propose suitable reversible jump operations to facilitate the challenging computations. We assess the performance of our method against benchmark examples, and use it to analyze a real world application with a large-scale climate model.

stat.ME↗

Comparative Study of Clustering Techniques for Real-Time Dynamic Model Reduction

Dynamic model reduction in power systems is necessary for improving computational efficiency. Traditional model reduction using linearized models or online analysis is not adequate to capture dynamic behaviors of the power system, especially with the new mix of intermittent generation and intelligent consumption making the power system more dynamic and non-linear. Real-time dynamic model reduction has emerged to fill this important need. This paper explores using clustering techniques to analyze real-time phasor measurements to identify groups of generators with similar behavior, as well as a representative generator from each group for dynamic model reduction. Two clustering techniques -- graph clustering and k-means -- are considered. These techniques are compared with a previously developed dynamic model reduction approach using Singular Value Decomposition. Two sample power grid data sets are used to test these different model reduction techniques. Based on the algorithms' relative performance, recommendations are provided for practical use.

physics.soc-ph↗

Regularity and spectral methods for two-sided fractional diffusion equations with a low-order term

We study regularity and numerical methods for two-sided fractional diffusion equations with a lower-order term. We show that the regularity of the solution in weighted Sobolev spaces can be greatly improved compared to that in standard Sobolev spaces. With this regularity, we improve higher-order convergence of a spectral Galerkin method. We present a spectral Petrov-Galerkin method and provide an optimal error estimate for the Petrov-Galerkin method. Numerical results are presented to verify our theoretical convergence orders.

math.NA↗

Uncertainty quantification of thermal conductivities from equilibrium molecular dynamics simulations

Equilibrium molecular dynamics (EMD) simulations along with the Green-Kubo formula have been widely used to calculate lattice thermal conductivities. Previous studies, however, focused primarily on the calculated thermal conductivities, with the uncertainty of the thermal conductivities remaining poorly understood. In this paper, we study the quantification of the uncertainty by using solid argon, silicon, and germanium as model material systems, and examine the origin of the observed uncertainty. We find that the uncertainty increases with the upper limit of the correlation time, $t_{\text{corre, UL}}$, and decreases with the total simulation time, $t_{\text{total}}$, whereas the velocity initialization seed, simulation domain size, temperature, and type of material have minimal effects. The relative uncertainties of the thermal conductivities for solid argon, silicon, and germanium under different simulation conditions all follow a similar trend, which can be fit with a "universal" square-root relation, as $σ_{k_x}/k_{x, \text{ave}} = 2(t_{\text{total}}/t_{\text{corre, UL}})^{-0.5}$. We have also conducted statistical analysis of the EMD-predicted thermal conductivities and derived a formula that correlates the relative error bound ($Q$), confidence level ($P$), $t_{\text{corre, UL}}$, $t_{\text{total}}$, and number of independent simulations ($N$). We recommend choosing $t_{\text{corre, UL}}$ to be 5-10 times the effective phonon relaxation time, $τ_{\text{eff}}$, and choosing $t_{\text{total}}$ and $N$ based on the desired relative error bound and confidence level. This study provides new insights into understanding the uncertainty of EMD-predicted thermal conductivities. It also provides a guideline for running EMD simulations to achieve a desired relative error bound with a desired confidence level and for reporting EMD-predicted thermal conductivities.

cond-mat.mtrl-sci↗

A two-level stochastic collocation method for semilinear elliptic equations with random coefficients

In this work, we propose a novel two-level discretization for solving semilinear elliptic equations with random coefficients. Motivated by the two-grid method for deterministic partial differential equations (PDEs) introduced by Xu \cite{xu1994novel}, our two-level stochastic collocation method utilizes a two-grid finite element discretization in the physical space and a two-level collocation method in the random domain. In particular, we solve semilinear equations on a coarse mesh $\mathcal{T}_H$ with a low level stochastic collocation (corresponding to the polynomial space $\mathcal{P}_{\boldsymbol{P}}$) and solve linearized equations on a fine mesh $\mathcal{T}_h$ using high level stochastic collocation (corresponding to the polynomial space $\mathcal{P}_{\boldsymbol{p}}$). We prove that the approximated solution obtained from this method achieves the same order of accuracy as that from solving the original semilinear problem directly by stochastic collocation method with $\mathcal{T}_h$ and $\mathcal{P}_{\boldsymbol{p}}$. The two-level method is computationally more efficient than the standard stochastic collocation method for solving nonlinear problems with random coefficients. Numerical experiments are provided to verify the theoretical results.

math.NA↗

A second-order difference scheme for the time fractional substantial diffusion equation

In this work, a second-order approximation of the fractional substantial derivative is presented by considering a modified shifted substantial Grünwald formula and its asymptotic expansion. Moreover, the proposed approximation is applied to a fractional diffusion equation with fractional substantial derivative in time. With the use of the fourth-order compact scheme in space, we give a fully discrete Grünwald-Letnikov-formula-based compact difference scheme and prove its stability and convergence by the energy method under smooth assumptions. In addition, the problem with nonsmooth solution is also discussed, and an improved algorithm is proposed to deal with the singularity of the fractional substantial derivative. Numerical examples show the reliability and efficiency of the scheme.

math.NA↗

Finite difference schemes for multi-term time-fractional mixed diffusion-wave equations

The multi-term time-fractional mixed diffusion-wave equations (TFMDWEs) are considered and the numerical method with its error analysis is presented in this paper. First, a $L2$ approximation is proved with first order accuracy to the Caputo fractional derivative of order $β\in (1,2).$ Then the approximation is applied to solve a one-dimensional TFMDWE and an implicit, compact difference scheme is constructed. Next, a rigorous error analysis of the proposed scheme is carried out by employing the energy method, and it is proved to be convergent with first order accuracy in time and fourth order in space, respectively. In addition, some results for the distributed order and two-dimensional extensions are also reported in this work. Subsequently, a practical fast solver with linearithmic complexity is provided with partial diagonalization technique. Finally, several numerical examples are given to demonstrate the accuracy and efficiency of proposed schemes.

math.NA↗

Generating Random Earthquake Events for PTHA

In order to perform probabilistic tsunami hazard assessment (PTHA) based on subduction zone earthquakes, it is necessary to start with a catalog of possible future events along with the annual probability of occurance, or a probability distribution of such events that can be easily sampled. For nearfield events, the distribution of slip on the fault can have a significant effect on the resulting tsunami. We present an approach to defining a probability distribution based on subdividing the fault geometry into many subfaults and prescribing a desired covariance matrix relating slip on one subfault to slip on any other subfault. The eigenvalues and eigenvectors of this matrix are then used to define a Karhunen-Loève expansion for random slip patterns. This is similar to a spectral representation of random slip based on Fourier series but conforms to a general fault geometry. We show that only a few terms in this series are needed to represent the features of the slip distribution that are most important in tsunami generation, first with a simple one-dimensional example where slip varies only in the down-dip direction and then on a portion of the Cascadia Subduction Zone.

math.NA↗

Enhancing Sparsity of Hermite Polynomial Expansions by Iterative Rotations

Compressive sensing has become a powerful addition to uncertainty quantification in recent years. This paper identifies new bases for random variables through linear mappings such that the representation of the quantity of interest is more sparse with new basis functions associated with the new random variables. This sparsity increases both the efficiency and accuracy of the compressive sensing-based uncertainty quantification method. Specifically, we consider rotation-based linear mappings which are determined iteratively for Hermite polynomial expansions. We demonstrate the effectiveness of the new method with applications in solving stochastic partial differential equations and high-dimensional ($\mathcal{O}(100)$) problems.

math.ST↗

A consistent hierarchy of generalized kinetic equation approximations to the chemical master equation applied to surface catalysis

We develop a hierarchy of approximations to the master equation for systems that exhibit translational invariance and finite-range spatial correlation. Each approximation within the hierarchy is a set of ordinary differential equations that considers spatial correlations of varying lattice distance; the assumption is that the full system will have finite spatial correlations and thus the behavior of the models within the hierarchy will approach that of the full system. We provide evidence of this convergence in the context of one- and two-dimensional numerical examples. Lower levels within the hierarchy that consider shorter spatial correlations, are shown to be up to three orders of magnitude faster than traditional kinetic Monte Carlo methods (KMC) for one-dimensional systems, while predicting similar system dynamics and steady states as KMC methods. We then test the hierarchy on a two-dimensional model for the oxidation of CO on RuO2(110), showing that low-order truncations of the hierarchy efficiently capture the essential system dynamics. By considering sequences of models in the hierarchy that account for longer spatial correlations, successive model predictions may be used to establish empirical approximation of error estimates. The hierarchy may be thought of as a class of generalized phenomenological kinetic models since each element of the hierarchy approximates the master equation and the lowest level in the hierarchy is identical to a simple existing phenomenological kinetic models.

physics.chem-ph↗

Gaussian process surrogates for failure detection: a Bayesian experimental design approach

An important task of uncertainty quantification is to identify {the probability of} undesired events, in particular, system failures, caused by various sources of uncertainties. In this work we consider the construction of Gaussian {process} surrogates for failure detection and failure probability estimation. In particular, we consider the situation that the underlying computer models are extremely expensive, and in this setting, determining the sampling points in the state space is of essential importance. We formulate the problem as an optimal experimental design for Bayesian inferences of the limit state (i.e., the failure boundary) and propose an efficient numerical scheme to solve the resulting optimization problem. In particular, the proposed limit-state inference method is capable of determining multiple sampling points at a time, and thus it is well suited for problems where multiple computer simulations can be performed in parallel. The accuracy and performance of the proposed method is demonstrated by both academic and practical examples.

stat.CO↗

Quantifying the influence of conformational uncertainty in biomolecular solvation

Biomolecules exhibit conformational fluctuations near equilibrium states, inducing uncertainty in various biological properties in a dynamic way. We have developed a general method to quantify the uncertainty of target properties induced by conformational fluctuations. Using a generalized polynomial chaos (gPC) expansion, we construct a surrogate model of the target property with respect to varying conformational states. We also propose a method to increase the sparsity of the gPC expansion by defining a set of conformational "active space" random variables. With the increased sparsity, we employ the compressive sensing method to accurately construct the surrogate model. We demonstrate the performance of the surrogate model by evaluating fluctuation-induced uncertainty in solvent-accessible surface area for the bovine trypsin inhibitor protein system and show that the new approach offers more accurate statistical information than standard Monte Carlo approaches. Further more, the constructed surrogate model also enables us to directly evaluate the target property under various conformational states, yielding a more accurate response surface than standard sparse grid collocation methods. In particular, the new method provides higher accuracy in high-dimensional systems, such as biomolecules, where sparse grid performance is limited by the accuracy of the computed quantity of interest. Our new framework is generalizable and can be used to investigate the uncertainty of a wide variety of target properties in biomolecular systems.

q-bio.BM↗

Parallel and Interacting Stochastic Approximation Annealing algorithms for global optimisation

We present the parallel and interacting stochastic approximation annealing (PISAA) algorithm, a stochastic simulation procedure for global optimisation, that extends and improves the stochastic approximation annealing (SAA) by using population Monte Carlo ideas. The standard SAA algorithm guarantees convergence to the global minimum when a square-root cooling schedule is used; however the efficiency of its performance depends crucially on its self-adjusting mechanism. Because its mechanism is based on information obtained from only a single chain, SAA may present slow convergence in complex optimisation problems. The proposed algorithm involves simulating a population of SAA chains that interact each other in a manner that ensures significant improvement of the self-adjusting mechanism and better exploration of the sampling space. Central to the proposed algorithm are the ideas of (i) recycling information from the whole population of Markov chains to design a more accurate/stable self-adjusting mechanism and (ii) incorporating more advanced proposals, such as crossover operations, for the exploration of the sampling space. PISAA presents a significantly improved performance in terms of convergence. PISAA can be implemented in parallel computing environments if available. We demonstrate the good performance of the proposed algorithm on challenging applications including Bayesian network learning and protein folding. Our numerical comparisons suggest that PISAA outperforms the simulated annealing, stochastic approximation annealing, and annealing evolutionary stochastic approximation Monte Carlo especially in high dimensional or rugged scenarios.

stat.CO↗

Classification of Spatio-Temporal Data via Asynchronous Sparse Sampling: Application to Flow Around a Cylinder

We present a novel method for the classification and reconstruction of time dependent, high-dimensional data using sparse measurements, and apply it to the flow around a cylinder. Assuming the data lies near a low dimensional manifold (low-rank dynamics) in space and has periodic time dependency with a sparse number of Fourier modes, we employ compressive sensing for accurately classifying the dynamical regime. We further show that we can reconstruct the full spatio-temporal behavior with these limited measurements, extending previous results of compressive sensing that apply for only a single snapshot of data. The method can be used for building improved reduced-order models and designing sampling/measurement strategies that leverage time asynchrony.

math.DS↗