SearcharxivSearch

arXiv subjects

Roger Ghanem

Publications and source records attributed to Roger Ghanem.

At least 19 recordsLinked to original sources

Uncertainty Quantification and Flow Dynamics in Rotating Detonation Engines

Rotating detonation engines (RDEs) are a critical technology for advancing combustion engines, particularly in applications requiring high efficiency and performance. Understanding the supersonic detonation structure and how various parameters influence these phenomena is essential for optimizing RDE design. In this study, we perform detailed simulations of detonations in an RDE and analyze how the flow patterns are affected by key parameters associated with the droplet arrangement within the engine. To further explore the system's sensitivity, we apply polynomial chaos expansion to investigate the propagation of uncertainties from input parameters to quantities of interest (QOIs). Additionally, we develop a framework to accurately characterize the joint distributions of QOIs with a limited number of simulations. Our findings indicate that the strategic release of droplets may be crucial for sustaining continuous detonation waves in the engine, and accurate representations of the solution (e..g via high-order chaos expansions) are essential to accurately capture the dependence between QOIs and input uncertainties. These insights provide a quantitative foundation for further optimization of RDE designs.

physics.flu-dyn

Stochastic Operator Learning for Chemistry in Non-Equilibrium Flows

This work presents a novel framework for physically consistent model error characterization and operator learning for reduced-order models of non-equilibrium chemical kinetics. By leveraging the Bayesian framework, we identify and infer sources of model and parametric uncertainty within the Coarse-Graining Methodology across a range of initial conditions. The model error is embedded into the chemical kinetics model to ensure that its propagation to quantities of interest remains physically consistent. For operator learning, we develop a methodology that separates time dynamics from other input parameters. Karhunen-Loeve Expansion (KLE) is employed to capture time dynamics, yielding temporal modes, while Polynomial Chaos Expansion (PCE) is subsequently used to map model error and input parameters to KLE coefficients. The proposed model offers three significant advantages: i) Separating time dynamics from other inputs ensures stability of chemistry surrogate when coupled with fluid solvers; ii) The framework fully accounts for model and parametric uncertainty, enabling robust probabilistic predictions; iii) The surrogate model is highly interpretable, with visualizable time modes and a PCE component that facilitates analytical calculation of sensitivity indices. We apply this framework to O2-O chemistry system under hypersonic flight conditions, validating it in both a 0D adiabatic reactor and coupled simulations with a fluid solver in a 1D shock case. Results demonstrate that the surrogate is stable during time integration, delivers physically consistent probabilistic predictions accounting for model and parametric uncertainty, and achieves maximum relative error below 10%. This work represents a significant step forward in enabling probabilistic predictions of non-equilibrium chemistry with coupled fluid solvers, offering a physically accurate approach for hypersonic flow predictions.

physics.comp-ph

Transient anisotropic kernel for probabilistic learning on manifolds

PLoM (Probabilistic Learning on Manifolds) is a method introduced in 2016 for handling small training datasets by projecting an It\^o equation from a stochastic dissipative Hamiltonian dynamical system, acting as the MCMC generator, for which the KDE-estimated probability measure with the training dataset is the invariant measure. PLoM performs a projection on a reduced-order vector basis related to the training dataset, using the diffusion maps (DMAPS) basis constructed with a time-independent isotropic kernel. In this paper, we propose a new ISDE projection vector basis built from a transient anisotropic kernel, providing an alternative to the DMAPS basis to improve statistical surrogates for stochastic manifolds with heterogeneous data. The construction ensures that for times near the initial time, the DMAPS basis coincides with the transient basis. For larger times, the differences between the two bases are characterized by the angle of their spanned vector subspaces. The optimal instant yielding the optimal transient basis is determined using an estimation of mutual information from Information Theory, which is normalized by the entropy estimation to account for the effects of the number of realizations used in the estimations. Consequently, this new vector basis better represents statistical dependencies in the learned probability measure for any dimension. Three applications with varying levels of statistical complexity and data heterogeneity validate the proposed theory, showing that the transient anisotropic kernel improves the learned probability measure.

stat.ML

Multifidelity uncertainty quantification with models based on dissimilar parameters

Multifidelity uncertainty quantification (MF UQ) sampling approaches have been shown to significantly reduce the variance of statistical estimators while preserving the bias of the highest-fidelity model, provided that the low-fidelity models are well correlated. However, maintaining a high level of correlation can be challenging, especially when models depend on different input uncertain parameters, which drastically reduces the correlation. Existing MF UQ approaches do not adequately address this issue. In this work, we propose a new sampling strategy that exploits a shared space to improve the correlation among models with dissimilar parametrization. We achieve this by transforming the original coordinates onto an auxiliary manifold using the adaptive basis (AB) method~\cite{Tipireddy2014}. The AB method has two main benefits: (1) it provides an effective tool to identify the low-dimensional manifold on which each model can be represented, and (2) it enables easy transformation of polynomial chaos representations from high- to low-dimensional spaces. This latter feature is used to identify a shared manifold among models without requiring additional evaluations. We present two algorithmic flavors of the new estimator to cover different analysis scenarios, including those with legacy and non-legacy high-fidelity data. We provide numerical results for analytical examples, a direct field acoustic test, and a finite element model of a nuclear fuel assembly. For all examples, we compare the proposed strategy against both single-fidelity and MF estimators based on the original model parametrization.

physics.data-an

Projection pursuit adaptation on polynomial chaos expansions

The present work addresses the issue of accurate stochastic approximations in high-dimensional parametric space using tools from uncertainty quantification (UQ). The basis adaptation method and its accelerated algorithm in polynomial chaos expansions (PCE) were recently proposed to construct low-dimensional approximations adapted to specific quantities of interest (QoI). The present paper addresses one difficulty with these adaptations, namely their reliance on quadrature point sampling, which limits the reusability of potentially expensive samples. Projection pursuit (PP) is a statistical tool to find the ``interesting'' projections in high-dimensional data and thus bypass the curse-of-dimensionality. In the present work, we combine the fundamental ideas of basis adaptation and projection pursuit regression (PPR) to propose a novel method to simultaneously learn the optimal low-dimensional spaces and PCE representation from given data. While this projection pursuit adaptation (PPA) can be entirely data-driven, the constructed approximation exhibits mean-square convergence to the solution of an underlying governing equation and is thus subject to the same physics constraints. The proposed approach is demonstrated on a borehole problem and a structural dynamics problem, demonstrating the versatility of the method and its ability to discover low-dimensional manifolds with high accuracy with limited data. In addition, the method can learn surrogate models for different quantities of interest while reusing the same data set.

math.NA

Probabilistic Learning on Manifolds (PLoM) with Partition

The probabilistic learning on manifolds (PLoM) introduced in 2016 has solved difficult supervised problems for the ``small data'' limit where the number N of points in the training set is small. Many extensions have since been proposed, making it possible to deal with increasingly complex cases. However, the performance limit has been observed and explained for applications for which $N$ is very small (50 for example) and for which the dimension of the diffusion-map basis is close to $N$. For these cases, we propose a novel extension based on the introduction of a partition in independent random vectors. We take advantage of this novel development to present improvements of the PLoM such as a simplified algorithm for constructing the diffusion-map basis and a new mathematical result for quantifying the concentration of the probability measure in terms of a probability upper bound. The analysis of the efficiency of this novel extension is presented through two applications.

stat.ME

Multi-market Oligopoly of Equal Capacity

We consider a variant of Cournot competition, where multiple firms allocate the same amount of resource across multiple markets. We prove that the game has a unique pure-strategy Nash equilibrium (NE), which is symmetric and is characterized by the maximal point of a "potential function". The NE is globally asymptotically stable under the gradient adjustment process, and is not socially optimal in general. An application is in transportation, where drivers allocate time over a street network.

math.OC

Data-based Discovery of Governing Equations

Most common mechanistic models are traditionally presented in mathematical forms to explain a given physical phenomenon. Machine learning algorithms, on the other hand, provide a mechanism to map the input data to output without explicitly describing the underlying physical process that generated the data. We propose a Data-based Physics Discovery (DPD) framework for automatic discovery of governing equations from observed data. Without a prior definition of the model structure, first a free-form of the equation is discovered, and then calibrated and validated against the available data. In addition to the observed data, the DPD framework can utilize available prior physical models, and domain expert feedback. When prior models are available, the DPD framework can discover an additive or multiplicative correction term represented symbolically. The correction term can be a function of the existing input variable to the prior model, or a newly introduced variable. In case a prior model is not available, the DPD framework discovers a new data-based standalone model governing the observations. We demonstrate the performance of the proposed framework on a real-world application in the aerospace industry.

cs.LG

Probabilistic learning on manifolds constrained by nonlinear partial differential equations for small datasets

A novel extension of the Probabilistic Learning on Manifolds (PLoM) is presented. It makes it possible to synthesize solutions to a wide range of nonlinear stochastic boundary value problems described by partial differential equations (PDEs) for which a stochastic computational model (SCM) is available and depends on a vector-valued random control parameter. The cost of a single numerical evaluation of this SCM is assumed to be such that only a limited number of points can be computed for constructing the training dataset (small data). Each point of the training dataset is made up realizations from a vector-valued stochastic process (the stochastic solution) and the associated random control parameter on which it depends. The presented PLoM constrained by PDE allows for generating a large number of learned realizations of the stochastic process and its corresponding random control parameter. These learned realizations are generated so as to minimize the vector-valued random residual of the PDE in the mean-square sense. Appropriate novel methods are developed to solve this challenging problem. Three applications are presented. The first one is a simple uncertain nonlinear dynamical system with a nonstationary stochastic excitation. The second one concerns the 2D nonlinear unsteady Navier-Stokes equations for incompressible flows in which the Reynolds number is the random control parameter. The last one deals with the nonlinear dynamics of a 3D elastic structure with uncertainties. The results obtained make it possible to validate the PLoM constrained by stochastic PDE but also provide further validation of the PLoM without constraint.

stat.ML

Drivers learn city-scale dynamic equilibrium

Understanding driver behavior in on-demand mobility services is crucial for designing efficient and sustainable transport models. Drivers' delivery strategy is well understood, but their search strategy and learning process still lack an empirically validated model. Here we provide a game-theoretic model of driver search strategy and learning dynamics, interpret the collective outcome in a thermodynamic framework, and verify its various implications empirically. We capture driver search strategies in a multi-market oligopoly model, which has a unique Nash equilibrium and is globally asymptotically stable. The equilibrium can therefore be obtained via heuristic learning rules where drivers pursue the incentive gradient or simply imitate others. To help understand city-scale phenomena, we offer a macroscopic view with the laws of thermodynamics. With 870 million trips of over 50k drivers in New York City, we show that the equilibrium well explains the spatiotemporal patterns of driver search behavior, and estimate an empirical constitutive relation. We find that new drivers learn the equilibrium within a year, and those who stay longer learn better. The collective response to new competition is also as predicted. Among empirical studies of driver strategy in on-demand services, our work examines the longest period, the most trips, and is the largest for taxi industry.

physics.soc-ph

Normal-bundle Bootstrap

Probabilistic models of data sets often exhibit salient geometric structure. Such a phenomenon is summed up in the manifold distribution hypothesis, and can be exploited in probabilistic learning. Here we present normal-bundle bootstrap (NBB), a method that generates new data which preserve the geometric structure of a given data set. Inspired by algorithms for manifold learning and concepts in differential geometry, our method decomposes the underlying probability measure into a marginalized measure on a learned data manifold and conditional measures on the normal spaces. The algorithm estimates the data manifold as a density ridge, and constructs new data by bootstrapping projection vectors and adding them to the ridge. We apply our method to the inference of density ridge and related statistics, and data augmentation to reduce overfitting.

stat.ML

Environmental Economics and Uncertainty: Review and a Machine Learning Outlook

Economic assessment in environmental science concerns the measurement or valuation of environmental impacts, adaptation, and vulnerability. Integrated assessment modeling is a unifying framework of environmental economics, which attempts to combine key elements of physical, ecological, and socioeconomic systems. Uncertainty characterization in integrated assessment varies by component models: uncertainties associated with mechanistic physical models are often assessed with an ensemble of simulations or Monte Carlo sampling, while uncertainties associated with impact models are evaluated by conjecture or econometric analysis. Manifold sampling is a machine learning technique that constructs a joint probability model of all relevant variables which may be concentrated on a low-dimensional geometric structure. Compared with traditional density estimation methods, manifold sampling is more efficient especially when the data is generated by a few latent variables. The manifold-constrained joint probability model helps answer policy-making questions from prediction, to response, and prevention. Manifold sampling is applied to assess risk of offshore drilling in the Gulf of Mexico.

econ.GN

Probabilistic Learning on Manifolds

This paper presents mathematical results in support of the methodology of the probabilistic learning on manifolds (PLoM) recently introduced by the authors, which has been used with success for analyzing complex engineering systems. The PLoM considers a given initial dataset constituted of a small number of points given in an Euclidean space, which are interpreted as independent realizations of a vector-valued random variable for which its non-Gaussian probability measure is unknown but is, \textit{a priori}, concentrated in an unknown subset of the Euclidean space. The objective is to construct a learned dataset constituted of additional realizations that allow the evaluation of converged statistics. A transport of the probability measure estimated with the initial dataset is done through a linear transformation constructed using a reduced-order diffusion-maps basis. In this paper, it is proven that this transported measure is a marginal distribution of the invariant measure of a reduced-order It\^o stochastic differential equation that corresponds to a dissipative Hamiltonian dynamical system. This construction allows for preserving the concentration of the probability measure. This property is shown by analyzing a distance between the random matrix constructed with the PLoM and the matrix representing the initial dataset, as a function of the dimension of the basis. It is further proven that this distance has a minimum for a dimension of the reduced-order diffusion-maps basis that is strictly smaller than the number of points in the initial dataset. Finally, a brief numerical application illustrates the mathematical results.

math.ST

Sampling of Bayesian posteriors with a non-Gaussian probabilistic learning on manifolds from a small dataset

This paper tackles the challenge presented by small-data to the task of Bayesian inference. A novel methodology, based on manifold learning and manifold sampling, is proposed for solving this computational statistics problem under the following assumptions: 1) neither the prior model nor the likelihood function are Gaussian and neither can be approximated by a Gaussian measure; 2) the number of functional input (system parameters) and functional output (quantity of interest) can be large; 3) the number of available realizations of the prior model is small, leading to the small-data challenge typically associated with expensive numerical simulations; the number of experimental realizations is also small; 4) the number of the posterior realizations required for decision is much larger than the available initial dataset. The method and its mathematical aspects are detailed. Three applications are presented for validation: The first two involve mathematical constructions aimed to develop intuition around the method and to explore its performance. The third example aims to demonstrate the operational value of the method using a more complex application related to the statistical inverse identification of the non-Gaussian matrix-valued random elasticity field of a damaged biological tissue (osteoporosis in a cortical bone) using ultrasonic waves.

stat.ML

Data-driven discovery of free-form governing differential equations

We present a method of discovering governing differential equations from data without the need to specify a priori the terms to appear in the equation. The input to our method is a dataset (or ensemble of datasets) corresponding to a particular solution (or ensemble of particular solutions) of a differential equation. The output is a human-readable differential equation with parameters calibrated to the individual particular solutions provided. The key to our method is to learn differentiable models of the data that subsequently serve as inputs to a genetic programming algorithm in which graphs specify computation over arbitrary compositions of functions, parameters, and (potentially differential) operators on functions. Differential operators are composed and evaluated using recursive application of automatic differentiation, allowing our algorithm to explore arbitrary compositions of operators without the need for human intervention. We also demonstrate an active learning process to identify and remedy deficiencies in the proposed governing equations.

cs.CE

Demand, Supply, and Performance of Street-Hail Taxi

Travel decisions are fundamental to understanding human mobility, urban economy, and sustainability, but measuring it is challenging and controversial. Previous studies of taxis are limited to taxi stands or hail markets at aggregate spatial units. Here we estimate the dynamic demand and supply of taxis in New York City (NYC) at street segment level, using in-vehicle Global Positioning System (GPS) data which preserve individual privacy. To this end, we model taxi demand and supply as non-stationary Poisson random fields on the road network, and pickups result from income-maximizing drivers searching for impatient passengers. With 868 million trip records of all 13,237 licensed taxis in NYC in 2009 - 2013, we show that while taxi demand are almost the same in 2011 and 2012, it declined about 2% in spring 2013, possibly caused by transportation network companies (TNCs) and fare raise. Contrary to common impression, street-hail taxis out-perform TNCs such as Uber in high-demand locations, suggesting a taxi/TNC regulation change to reduce congestion and pollution. We show that our demand estimates are stable at different supply levels and across years, a property not observed in existing matching functions. We also validate that taxi pickups can be modeled as Poisson processes. Our method is thus simple, feasible, and reliable in estimating street-hail taxi activities at a high spatial resolution; it helps quantify the ongoing discussion on congestion charges to taxis and TNCs.

physics.soc-ph

Reduced Wiener Chaos representation of random fields via basis adaptation and projection

A new characterization of random fields appearing in physical models is presented that is based on their well-known Homogeneous Chaos expansions. We take advantage of the adaptation capabilities of these expansions where the core idea is to rotate the basis of the underlying Gaussian Hilbert space, in order to achieve reduced functional representations that concentrate the induced probability measure in a lower dimensional subspace. For a smooth family of rotations along the domain of interest, the uncorrelated Gaussian inputs are transformed into a Gaussian process, thus introducing a mesoscale that captures intermediate characteristics of the quantity of interest.

stat.ME

Data-driven probability concentration and sampling on manifold

A new methodology is proposed for generating realizations of a random vector with values in a finite-dimensional Euclidean space that are statistically consistent with a data set of observations of this vector. The probability distribution of this random vector, while a-priori not known, is presumed to be concentrated on an unknown subset of the Euclidean space. A random matrix is introduced whose columns are independent copies of the random vector and for which the number of columns is the number of data points in the data set. The approach is based on the use of (i) the multidimensional kernel-density estimation method for estimating the probability distribution of the random matrix, (ii) a MCMC method for generating realizations for the random matrix, (iii) the diffusion-maps approach for discovering and characterizing the geometry and the structure of the data set, and (iv) a reduced-order representation of the random matrix, which is constructed using the diffusion-maps vectors associated with the first eigenvalues of the transition matrix relative to the given data set. The convergence aspects of the proposed methodology are analyzed and a numerical validation is explored through three applications of increasing complexity. The proposed method is found to be robust to noise levels and data complexity as well as to the intrinsic dimension of data and the size of experimental data sets. Both the methodology and the underlying mathematical framework presented in this paper contribute new capabilities and perspectives at the interface of uncertainty quantification, statistical data analysis, stochastic modeling and associated statistical inverse problems.

math.PR