SearcharxivSearch

arXiv subjects

Stefano Marelli

Publications and source records attributed to Stefano Marelli.

At least 19 recordsLinked to original sources

Time-variant reliability using time-dependent surrogate models

Time-variant reliability analysis is a critical task for ensuring the safety of engineering dynamical systems subjected to stochastic excitations. However, assessing failure probability for realistic systems with Monte-Carlo simulation-based methods is often computationally intractable due to the high cost of the underlying models and the large number of simulations required. While surrogate models such as polynomial chaos expansions or Kriging are well-established for time-invariant reliability problems, their direct application to time-dependent systems remains challenging. This chapter introduces two advanced surrogate modeling frameworks designed specifically for dynamical systems: manifold-NARX (mNARX) and functional NARX (F-NARX). The mNARX approach constructs the surrogate on a reduced-order manifold of auxiliary state variables, enabling the efficient handling of high-dimensional inputs by embedding physical insight into a regression formulation. Conversely, the F-NARX framework exploits the functional nature of system trajectories, extracting principal component features from continuous time windows to mitigate issues associated with discrete lag selection and long-memory effects. We demonstrate the efficacy of these methods on two benchmark reliability problems: a stochastic quarter-car model and a hysteretic Bouc-Wen oscillator. The results highlight that, when combined with suitably biased experimental designs, both frameworks accurately capture the tail behavior of the system response, enabling precise and efficient estimation of first-passage probabilities.

stat.ME

Probabilistic function-on-function nonlinear autoregressive model for emulation and reliability analysis of stochastic dynamical systems

Constructing accurate and computationally efficient surrogate models (or emulators) for predicting dynamical system responses is critical in many engineering domains, yet remains challenging due to the strongly nonlinear and high-dimensional mapping from external excitations and system parameters to system responses. This work introduces a novel Function-on-Function Nonlinear AutoRegressive model with eXogenous inputs (F2NARX), which reformulates the recently proposed $\mathcal{F}$-NARX method from a function-on-function regression perspective. The proposed framework substantially improves predictive efficiency while maintaining high accuracy. By combining principal component analysis with Gaussian process regression, F2NARX further enables probabilistic predictions of dynamical responses via the unscented transform in an autoregressive manner. Such probabilistic prediction capabilities further facilitate active learning for first-passage probability evaluation. The effectiveness of the method is demonstrated through case studies of varying complexity. Results show that F2NARX outperforms state-of-the-art NARX model by orders of magnitude in efficiency while achieving higher accuracy in general. Meanwhile, the active learning approach enables accurate estimation of first-passage failure probabilities for dynamical systems using only a small number of training time histories.

math.DS

Surrogate-Accelerated Bayesian Inversion for Exoplanet Interior Characterization

Characterizing the interior structure of exoplanets is an inverse problem often solved using Bayesian inference, but this approach is hampered by the high computational cost of planetary structure models. To overcome this barrier, we present a robust framework that accelerates inference by replacing the computationally expensive physics-based forward model with a fast polynomial chaos-Kriging (PCK) surrogate directly within a Markov chain Monte Carlo (MCMC) sampling loop. We rigorously validate our approach using a suite of tests, including a direct comparison against a benchmark MCMC inference using the full forward model, and a large-scale coverage study with 1000 synthetic test cases to demonstrate the statistical reliability of our inferred credible intervals. Our surrogate-assisted framework achieves a computational speedup of over 2 orders of magnitude (factor of $\sim$320), reducing single-CPU inference times from days to minutes. This efficiency is achieved with a surrogate that requires only a few hundred forward model evaluations for training \rev{for a single planet}. This data efficiency provides significant flexibility for model developments and a clear advantage over common machine learning approaches, which typically demand vast training sets ($>10^6$ model runs) and intensive pre-computation. The PCK surrogate maintains high fidelity with $R^2 > 0.99$ for most scenarios, and root-mean-square errors typically an order of magnitude smaller than observational uncertainties. This efficiency enables large scale population studies while preserving statistical robustness, which is computationally impractical with traditional methods.

astro-ph.EP

Bayesian full waveform inversion with sequential surrogate model refinement

Bayesian formulations of inverse problems are attractive for their ability to incorporate prior knowledge and update probabilistic models as new data become available. Markov chain Monte Carlo (MCMC) methods sample posterior probability density functions (pdfs) but require accurate prior models and many likelihood evaluations. Dimensionality-reduction methods, such as principal component analysis (PCA), can help define the prior and train surrogate models that efficiently approximate costly forward solvers. However, for problems like full waveform inversion, the complex input/output relations often cannot be captured well by surrogate models trained only on prior samples, leading to biased results. Including samples from high-posterior-probability regions can improve accuracy, but these regions are hard to identify in advance. We propose an iterative method that progressively refines the surrogate model. Starting with low-frequency data, we train an initial surrogate and perform an MCMC inversion. The resulting posterior samples are then used to retrain the surrogate, allowing us to expand the frequency bandwidth in the next inversion step. Repeating this process reduces model errors and improves the surrogate's accuracy over the relevant input domain. Ultimately, we obtain a highly accurate surrogate across the full bandwidth, enabling a final MCMC inversion. Numerical results from 2D synthetic crosshole Ground Penetrating Radar (GPR) examples show that our method outperforms ray-based approaches and those relying solely on prior sampling. The overall computational cost is reduced by about two orders of magnitude compared to full finite-difference time-domain modeling.

physics.geo-ph

Reliability analysis for non-deterministic limit-states using stochastic emulators

Reliability analysis is a sub-field of uncertainty quantification that assesses the probability of a system performing as intended under various uncertainties. Traditionally, this analysis relies on deterministic models, where experiments are repeatable, i.e., they produce consistent outputs for a given set of inputs. However, real-world systems often exhibit stochastic behavior, leading to non-repeatable outcomes. These so-called stochastic simulators produce different outputs each time the model is run, even with fixed inputs. This paper formally introduces reliability analysis for stochastic models and addresses it by using suitable surrogate models to lower its typically high computational cost. Specifically, we focus on the recently introduced generalized lambda models and stochastic polynomial chaos expansions. These emulators are designed to learn the inherent randomness of the simulator's response and enable efficient uncertainty quantification at a much lower cost than traditional Monte Carlo simulation. We validate our methodology through three case studies. First, using an analytical function with a closed-form solution, we demonstrate that the emulators converge to the correct solution. Second, we present results obtained from the surrogates using a toy example of a simply supported beam. Finally, we apply the emulators to perform reliability analysis on a realistic wind turbine case study, where only a dataset of simulation results is available.

stat.CO

Surrogate modeling with functional nonlinear autoregressive models (F-NARX)

We propose a novel functional approach to surrogate modeling of dynamical systems with exogenous inputs. This approach, named Functional Nonlinear AutoRegressive with eXogenous inputs (F-NARX), approximates the system response based on temporal features of the exogenous inputs and the system response. This marks a major step away from the discrete-time-centric approach of classical NARX models, which determines the relationship between selected time steps of the input/output time series. By modeling the system in a time-feature space, F-NARX takes advantage of the temporal smoothness of the process being modeled, providing more stable predictions and reducing the dependence of model performance on the discretization of the time axis. In this work, we introduce an F-NARX implementation based on principal component analysis and polynomial regression. To further improve prediction accuracy, we also introduce a modified hybrid least angle regression approach to identify a sparse model structure and minimize the expected forecast error, rather than the one-step-ahead prediction error. We investigate the behavior and capabilities of our F-NARX implementation on two case studies: an eight-story building under wind loading and a three-story steel frame under seismic loading. Our results demonstrate that F-NARX has several favorable properties that make it well-suited to surrogate modeling applications.

stat.ME

Reliability analysis for data-driven noisy models using active learning

Reliability analysis aims at estimating the failure probability of an engineering system. It often requires multiple runs of a limit-state function, which usually relies on computationally intensive simulations. Traditionally, these simulations have been considered deterministic, i.e., running them multiple times for a given set of input parameters always produces the same output. However, this assumption does not always hold, as many studies in the literature report non-deterministic computational simulations (also known as noisy models). In such cases, running the simulations multiple times with the same input will result in different outputs. Similarly, data-driven models that rely on real-world data may also be affected by noise. This characteristic poses a challenge when performing reliability analysis, as many classical methods, such as FORM and SORM, are tailored to deterministic models. To bridge this gap, this paper provides a novel methodology to perform reliability analysis on models contaminated by noise. In such cases, noise introduces latent uncertainty into the reliability estimator, leading to an incorrect estimation of the real underlying reliability index, even when using Monte Carlo simulation. To overcome this challenge, we propose the use of denoising regression-based surrogate models within an active learning reliability analysis framework. Specifically, we combine Gaussian process regression with a noise-aware learning function to efficiently estimate the probability of failure of the underlying noise-free model. We showcase the effectiveness of this methodology on standard benchmark functions and a finite element model of a realistic structural frame.

stat.CO

Uncertainty-aware multi-fidelity surrogate modeling with noisy data

Emulating high-accuracy computationally expensive models is crucial for tasks requiring numerous model evaluations, such as uncertainty quantification and optimization. When lower-fidelity models are available, they can be used to improve the predictions of high-fidelity models. Multi-fidelity surrogate models combine information from sources of varying fidelities to construct an efficient surrogate model. However, in real-world applications, uncertainty is present in both high- and low-fidelity models due to measurement or numerical noise, as well as lack of knowledge due to the limited experimental design budget. This paper introduces a comprehensive framework for multi-fidelity surrogate modeling that handles noise-contaminated data and is able to estimate the underlying noise-free high-fidelity model. Our methodology quantitatively incorporates the different types of uncertainty affecting the problem and emphasizes on delivering precise estimates of the uncertainty in its predictions both with respect to the underlying high-fidelity model and unseen noise-contaminated high-fidelity observations, presented through confidence and prediction intervals, respectively. Additionally, the proposed framework offers a natural approach to combining physical experiments and computational models by treating noisy experimental data as high-fidelity sources and white-box computational models as their low-fidelity counterparts. The effectiveness of our methodology is showcased through synthetic examples and a wind turbine application.

stat.ME

Comparison of Probabilistic Structural Reliability Methods for Ultimate Limit State Assessment of Wind Turbines

The probabilistic design of offshore wind turbines aims to ensure structural safety in a cost-effective way. This involves conducting structural reliability assessments for different design options and considering different structural responses. There are several structural reliability methods, and this paper will apply and compare different approaches in some simplified case studies. In particular, the well known environmental contour method will be compared to a more novel approach based on sequential sampling and Gaussian processes regression for an ultimate limit state case study. For one of the case studies, results will also be compared to results from a brute force simulation approach. Interestingly, the comparison is very different from the two case studies. In one of the cases the environmental contours method agrees well with the sequential sampling method but in the other, results vary considerably. Probably, this can be explained by the violation of some of the assumptions associated with the environmental contour approach, i.e. that the short-term variability of the response is large compared to the long-term variability of the environmental conditions. Results from this simple comparison study suggests that the sequential sampling method can be a robust and computationally effective approach for structural reliability assessment.

stat.AP

Bayesian tomography using polynomial chaos expansion and deep generative networks

Implementations of Markov chain Monte Carlo (MCMC) methods need to confront two fundamental challenges: accurate representation of prior information and efficient evaluation of likelihoods. Principal component analysis (PCA) and related techniques can in some cases facilitate the definition and sampling of the prior distribution, as well as the training of accurate surrogate models, using for instance, polynomial chaos expansion (PCE). However, complex geological priors with sharp contrasts necessitate more complex dimensionality-reduction techniques, such as, deep generative models (DGMs). By sampling a low-dimensional prior probability distribution defined in the low-dimensional latent space of such a model, it becomes possible to efficiently sample the physical domain at the price of a generator that is typically highly non-linear. Training a surrogate that is capable of capturing intricate non-linear relationships between latent parameters and outputs of forward modeling presents a notable challenge. Indeed, while PCE models provide high accuracy when the input-output relationship can be effectively approximated by relatively low-degree multivariate polynomials, this condition is typically not met when employing latent variables derived from DGMs. In this contribution, we present a strategy combining the excellent reconstruction performances of a variational autoencoder (VAE) with the accuracy of PCA-PCE surrogate modeling in the context of Bayesian ground penetrating radar (GPR) traveltime tomography. Within the MCMC process, the parametrization of the VAE is leveraged for prior exploration and sample proposals. Concurrently, surrogate modeling is conducted using PCE, which operates on either globally or locally defined principal components of the VAE samples under examination.

physics.geo-ph

Emulating the dynamics of complex systems using autoregressive models on manifolds (mNARX)

We propose a novel surrogate modelling approach to efficiently and accurately approximate the response of complex dynamical systems driven by time-varying exogenous excitations over extended time periods. Our approach, namely manifold nonlinear autoregressive modelling with exogenous input (mNARX), involves constructing a problem-specific exogenous input manifold that is optimal for constructing autoregressive surrogates. The manifold, which forms the core of mNARX, is constructed incrementally by incorporating the physics of the system, as well as prior expert- and domain- knowledge. Because mNARX decomposes the full problem into a series of smaller sub-problems, each with a lower complexity than the original, it scales well with the complexity of the problem, both in terms of training and evaluation costs of the final surrogate. Furthermore, mNARX synergizes well with traditional dimensionality reduction techniques, making it highly suitable for modelling dynamical systems with high-dimensional exogenous inputs, a class of problems that is typically challenging to solve. Since domain knowledge is particularly abundant in physical systems, such as those found in civil and mechanical engineering, mNARX is well suited for these applications. We demonstrate that mNARX outperforms traditional autoregressive surrogates in predicting the response of a classical coupled spring-mass system excited by a one-dimensional random excitation. Additionally, we show that mNARX is well suited for emulating very high-dimensional time- and state-dependent systems, even when affected by active controllers, by surrogating the dynamics of a realistic aero-servo-elastic onshore wind turbine simulator. In general, our results demonstrate that mNARX offers promising prospects for modelling complex dynamical systems, in terms of accuracy and efficiency.

stat.CO

Global sensitivity analysis using derivative-based sparse Poincaré chaos expansions

Variance-based global sensitivity analysis, in particular Sobol' analysis, is widely used for determining the importance of input variables to a computational model. Sobol' indices can be computed cheaply based on spectral methods like polynomial chaos expansions (PCE). Another choice are the recently developed Poincaré chaos expansions (PoinCE), whose orthonormal tensor-product basis is generated from the eigenfunctions of one-dimensional Poincaré differential operators. In this paper, we show that the Poincaré basis is the unique orthonormal basis with the property that partial derivatives of the basis form again an orthogonal basis with respect to the same measure as the original basis. This special property makes PoinCE ideally suited for incorporating derivative information into the surrogate modelling process. Assuming that partial derivative evaluations of the computational model are available, we compute spectral expansions in terms of Poincaré basis functions or basis partial derivatives, respectively, by sparse regression. We show on two numerical examples that the derivative-based expansions provide accurate estimates for Sobol' indices, even outperforming PCE in terms of bias and variance. In addition, we derive an analytical expression based on the PoinCE coefficients for a second popular sensitivity index, the derivative-based sensitivity measure (DGSM), and explore its performance as upper bound to the corresponding total Sobol' indices.

stat.CO

Reliability analysis of arbitrary systems based on active learning and global sensitivity analysis

System reliability analysis aims at computing the probability of failure of an engineering system given a set of uncertain inputs and limit state functions. Active-learning solution schemes have been shown to be a viable tool but as of yet they are not as efficient as in the context of component reliability analysis. This is due to some peculiarities of system problems, such as the presence of multiple failure modes and their uneven contribution to failure, or the dependence on the system configuration (e.g., series or parallel). In this work, we propose a novel active learning strategy designed for solving general system reliability problems. This algorithm combines subset simulation and Kriging/PC-Kriging, and relies on an enrichment scheme tailored to specifically address the weaknesses of this class of methods. More specifically, it relies on three components: (i) a new learning function that does not require the specification of the system configuration, (ii) a density-based clustering technique that allows one to automatically detect the different failure modes, and (iii) sensitivity analysis to estimate the contribution of each limit state to system failure so as to select only the most relevant ones for enrichment. The proposed method is validated on two analytical examples and compared against results gathered in the literature. Finally, a complex engineering problem related to power transmission is solved, thereby showcasing the efficiency of the proposed method in a real-case scenario.

stat.ME

A spectral surrogate model for stochastic simulators computed from trajectory samples

Stochastic simulators are non-deterministic computer models which provide a different response each time they are run, even when the input parameters are held at fixed values. They arise when additional sources of uncertainty are affecting the computer model, which are not explicitly modeled as input parameters. The uncertainty analysis of stochastic simulators requires their repeated evaluation for different values of the input variables, as well as for different realizations of the underlying latent stochasticity. The computational cost of such analyses can be considerable, which motivates the construction of surrogate models that can approximate the original model and its stochastic response, but can be evaluated at much lower cost. We propose a surrogate model for stochastic simulators based on spectral expansions. Considering a certain class of stochastic simulators that can be repeatedly evaluated for the same underlying random event, we view the simulator as a random field indexed by the input parameter space. For a fixed realization of the latent stochasticity, the response of the simulator is a deterministic function, called trajectory. Based on samples from several such trajectories, we approximate the latter by sparse polynomial chaos expansion and compute analytically an extended Karhunen-Loève expansion (KLE) to reduce its dimensionality. The uncorrelated but dependent random variables of the KLE are modeled by advanced statistical techniques such as parametric inference, vine copula modeling, and kernel density estimation. The resulting surrogate model approximates the marginals and the covariance function, and allows to obtain new realizations at low computational cost. We observe that in our numerical examples, the first mode of the KLE is by far the most important, and investigate this phenomenon and its implications.

stat.CO

Bayesian tomography with prior-knowledge-based parametrization and surrogate modeling

We present a Bayesian tomography framework operating with prior-knowledge-based parametrization that is accelerated by surrogate models. Standard high-fidelity forward solvers solve wave equations with natural spatial parametrizations based on fine discretization. Similar parametrizations, typically involving tens of thousand of variables, are usually employed to parameterize the subsurface in tomography applications. When the data do not allow to resolve details at such finely parameterized scales, it is often beneficial to instead rely on a prior-knowledge-based parametrization defined on a lower dimension domain (or manifold). Due to the increased identifiability in the reduced domain, the concomitant inversion is better constrained and generally faster. We illustrate the potential of a prior-knowledge-based approach by considering ground penetrating radar (GPR) travel-time tomography in a crosshole configuration. An effective parametrization of the input (i.e., the permittivity distributions) and compression in the output (i.e., the travel-time gathers) spaces are achieved via data-driven principal component decomposition based on random realizations of the prior Gaussian-process model with a truncation determined by the performances of the standard solver on the full and reduced model domains. To accelerate the inversion process, we employ a high-fidelity polynomial chaos expansion (PCE) surrogate model. We show that a few hundreds design data sets is sufficient to provide reliable Markov chain Monte Carlo inversion. Appropriate uncertainty quantification is achieved by reintroducing the truncated higher-order principle components in the original model space after inversion on the manifold and by adapting a likelihood function that accounts for the fact that the truncated higher-order components are not completely located in the null-space.

physics.geo-ph

Automatic selection of basis-adaptive sparse polynomial chaos expansions for engineering applications

Sparse polynomial chaos expansions (PCE) are an efficient and widely used surrogate modeling method in uncertainty quantification for engineering problems with computationally expensive models. To make use of the available information in the most efficient way, several approaches for so-called basis-adaptive sparse PCE have been proposed to determine the set of polynomial regressors ("basis") for PCE adaptively. The goal of this paper is to help practitioners identify the most suitable methods for constructing a surrogate PCE for their model. We describe three state-of-the-art basis-adaptive approaches from the recent sparse PCE literature and conduct an extensive benchmark in terms of global approximation accuracy on a large set of computational models. Investigating the synergies between sparse regression solvers and basis adaptivity schemes, we find that the choice of the proper solver and basis-adaptive scheme is very important, as it can result in more than one order of magnitude difference in performance. No single method significantly outperforms the others, but dividing the analysis into classes (regarding input dimension and experimental design size), we are able to identify specific sparse solver and basis adaptivity combinations for each class that show comparatively good performance. To further improve on these findings, we introduce a novel solver and basis adaptivity selection scheme guided by cross-validation error. We demonstrate that this automatic selection procedure provides close-to-optimal results in terms of accuracy, and significantly more robust solutions, while being more general than the case-by-case recommendations obtained by the benchmark.

stat.CO

Sparse Polynomial Chaos Expansions: Literature Survey and Benchmark

Sparse polynomial chaos expansions (PCE) are a popular surrogate modelling method that takes advantage of the properties of PCE, the sparsity-of-effects principle, and powerful sparse regression solvers to approximate computer models with many input parameters, relying on only few model evaluations. Within the last decade, a large number of algorithms for the computation of sparse PCE have been published in the applied math and engineering literature. We present an extensive review of the existing methods and develop a framework for classifying the algorithms. Furthermore, we conduct a unique benchmark on a selection of methods to identify which approaches work best in practical applications. Comparing their accuracy on several benchmark models of varying dimensionality and complexity, we find that the choice of sparse regression solver and sampling scheme for the computation of a sparse PCE surrogate can make a significant difference, of up to several orders of magnitude in the resulting mean-squared error. Different methods seem to be superior in different regimes of model dimensionality and experimental design size.

math.NA

Machine learning applied to simulations of collisions between rotating, differentiated planets

In the late stages of terrestrial planet formation, pairwise collisions between planetary-sized bodies act as the fundamental agent of planet growth. These collisions can lead to either growth or disruption of the bodies involved and are largely responsible for shaping the final characteristics of the planets. Despite their critical role in planet formation, an accurate treatment of collisions has yet to be realized. While semi-analytic methods have been proposed, they remain limited to a narrow set of post-impact properties and have only achieved relatively low accuracies. However, the rise of machine learning and access to increased computing power have enabled novel data-driven approaches. In this work, we show that data-driven emulation techniques are capable of predicting the outcome of collisions with high accuracy and are generalizable to any quantifiable post-impact quantity. In particular, we focus on the dataset requirements, training pipeline, and regression performance for four distinct data-driven techniques from machine learning (ensemble methods and neural networks) and uncertainty quantification (Gaussian processes and polynomial chaos expansion). We compare these methods to existing analytic and semi-analytic methods. Such data-driven emulators are poised to replace the methods currently used in N-body simulations. This work is based on a new set of 10,700 SPH simulations of pairwise collisions between rotating, differentiated bodies at all possible mutual orientations.

astro-ph.EP