SearcharxivSearch

arXiv subjects

Alexandre Tartakovsky

Publications and source records attributed to Alexandre Tartakovsky.

At least 19 recordsLinked to original sources

A Trainable-by-Parts Operator Learning Framework: Bridging DeepONet and Karhunen-Loeve Expansions for Large-Scale Applications

Training operator-learning models for large-scale problems governed by partial differential equations (PDEs) is challenging due to the curse of dimensionality, memory constraints, and limited training data. These challenges arise in many scientific and engineering applications, including subsurface flow, climate modeling, and geological carbon storage (GCS). In this work, we propose a scalable operator-learning framework based on the Karhunen-Loeve Deep Neural Network (KL-DNN) and demonstrate its performance for modeling GCS. The model is trained on a dataset comprising 100 samples of large-scale simulations in a three-dimensional domain with 1.7 million cells and 50 time steps. The KL-DNN method constructs latent spaces using low-rank singular value decomposition of static properties and a nested Karhunen-Loeve expansion for dynamic pressure fields, enabling full-resolution predictions without subsampling or spatial coarsening. The KL-DNN model achieves an average root mean square error (RMSE) of 1.1 psi for pressure (0.04% relative error with respect to the average pressure in the domain) and RMSE of 0.0146 for CO2 saturation (5% relative error with respect to the average saturation inside the plume). The model requires 20 minutes of training on a single GPU, representing a 19% reduction in the pressure errors, 7% reduction in the saturation error, and a two-order-of-magnitude speedup compared to DeepONet trained on the same dataset. These results, along with inference time of less than one minute, establish the proposed model as a practical and accurate solution for large-scale PDE problems, enabling rapid uncertainty quantification, history matching, and real-time decision support.

cs.LG

Physics-Informed Neural Networks and Radial Basis Functions for PDEs with Dirac Delta Sources

Physics-Informed Neural Networks (PINNs) are a machine learning method for solving forward and inverse Partial Differential Equations (PDEs). When applied to PDEs with Dirac delta functions in the forcing terms, boundary conditions, or initial conditions, PINNs require approximating them with smooth surrogate functions, a practice that can introduce significant modeling errors. In this work, we exploit the interpretation of PINNs as Residual Least Squares (RLS) methods and show that this perspective enables direct treatment of Dirac delta terms by integrating the weak-form equation. Among RLS formulations other than PINN, we focus on the Radial Basis Function (RBF) expansion (also known as a single-layer RBF Network). We show that while integrating out the Dirac delta in PINNs causes residuals to fail to converge to zero, RBF-RLS consistently provides good forward and inverse solutions to transport problems. We explain this finding using the Neural Tangent Kernel (NTK) theory. We test both approaches on linear PDEs that represent groundwater flow and transport in porous media and rivers. We solve inverse problems to fit synthetic data, noisy synthetic data, and real-world measurements.

cs.LG

Mathematics of Digital Twins and Transfer Learning for PDE Models

We define a digital twin (DT) of a physical system governed by partial differential equations (PDEs) as a model for real-time simulations and control of the system behavior under changing conditions. We construct DTs using the Karhunen-Lo\`{e}ve Neural Network (KL-NN) surrogate model and transfer learning (TL). The surrogate model allows fast inference and differentiability with respect to control parameters for control and optimization. TL is used to retrain the model for new conditions with minimal additional data. We employ the moment equations to analyze TL and identify parameters that can be transferred to new conditions. The proposed analysis also guides the control variable selection in DT to facilitate efficient TL. For linear PDE problems, the non-transferable parameters in the KL-NN surrogate model can be exactly estimated from a single solution of the PDE corresponding to the mean values of the control variables under new target conditions. Retraining an ML model with a single solution sample is known as one-shot learning, and our analysis shows that the one-shot TL is exact for linear PDEs. For nonlinear PDE problems, transferring of any parameters introduces errors. For a nonlinear diffusion PDE model, we find that for a relatively small range of control variables, some surrogate model parameters can be transferred without introducing a significant error, some can be approximately estimated from the mean-field equation, and the rest can be found using a linear residual least square problem or an ordinary linear least square problem if a small labeled dataset for new conditions is available. The former approach results in a one-shot TL while the latter approach is an example of a few-shot TL. Both methods are approximate for the nonlinear PDEs.

cs.LG

Differentiable modeling to unify machine learning and physical models and advance Geosciences

Process-Based Modeling (PBM) and Machine Learning (ML) are often perceived as distinct paradigms in the geosciences. Here we present differentiable geoscientific modeling as a powerful pathway toward dissolving the perceived barrier between them and ushering in a paradigm shift. For decades, PBM offered benefits in interpretability and physical consistency but struggled to efficiently leverage large datasets. ML methods, especially deep networks, presented strong predictive skills yet lacked the ability to answer specific scientific questions. While various methods have been proposed for ML-physics integration, an important underlying theme -- differentiable modeling -- is not sufficiently recognized. Here we outline the concepts, applicability, and significance of differentiable geoscientific modeling (DG). "Differentiable" refers to accurately and efficiently calculating gradients with respect to model variables, critically enabling the learning of high-dimensional unknown relationships. DG refers to a range of methods connecting varying amounts of prior knowledge to neural networks and training them together, capturing a different scope than physics-guided machine learning and emphasizing first principles. Preliminary evidence suggests DG offers better interpretability and causality than ML, improved generalizability and extrapolation capability, and strong potential for knowledge discovery, while approaching the performance of purely data-driven ML. DG models require less training data while scaling favorably in performance and efficiency with increasing amounts of data. With DG, geoscientists may be better able to frame and investigate questions, test hypotheses, and discover unrecognized linkages.

cs.LG

Enhanced physics-constrained deep neural networks for modeling vanadium redox flow battery

Numerical modeling and simulation have become indispensable tools for advancing a comprehensive understanding of the underlying mechanisms and cost-effective process optimization and control of flow batteries. In this study, we propose an enhanced version of the physics-constrained deep neural network (PCDNN) approach [1] to provide high-accuracy voltage predictions in the vanadium redox flow batteries (VRFBs). The purpose of the PCDNN approach is to enforce the physics-based zero-dimensional (0D) VRFB model in a neural network to assure model generalization for various battery operation conditions. Limited by the simplifications of the 0D model, the PCDNN cannot capture sharp voltage changes in the extreme SOC regions. To improve the accuracy of voltage prediction at extreme ranges, we introduce a second (enhanced) DNN to mitigate the prediction errors carried from the 0D model itself and call the resulting approach enhanced PCDNN (ePCDNN). By comparing the model prediction with experimental data, we demonstrate that the ePCDNN approach can accurately capture the voltage response throughout the charge--discharge cycle, including the tail region of the voltage discharge curve. Compared to the standard PCDNN, the prediction accuracy of the ePCDNN is significantly improved. The loss function for training the ePCDNN is designed to be flexible by adjusting the weights of the physics-constrained DNN and the enhanced DNN. This allows the ePCDNN framework to be transferable to battery systems with variable physical model fidelity.

physics.chem-ph

Machine Learning in Heterogeneous Porous Materials

The "Workshop on Machine learning in heterogeneous porous materials" brought together international scientific communities of applied mathematics, porous media, and material sciences with experts in the areas of heterogeneous materials, machine learning (ML) and applied mathematics to identify how ML can advance materials research. Within the scope of ML and materials research, the goal of the workshop was to discuss the state-of-the-art in each community, promote crosstalk and accelerate multi-disciplinary collaborative research, and identify challenges and opportunities. As the end result, four topic areas were identified: ML in predicting materials properties, and discovery and design of novel materials, ML in porous and fractured media and time-dependent phenomena, Multi-scale modeling in heterogeneous porous materials via ML, and Discovery of materials constitutive laws and new governing equations. This workshop was part of the AmeriMech Symposium series sponsored by the National Academies of Sciences, Engineering and Medicine and the U.S. National Committee on Theoretical and Applied Mechanics.

cs.LG

Autonomous Inversion of In Situ Deformation Measurement Data for CO2 Storage Decision Support

Current methods of estimating the change in stress caused by injecting fluid into subsurface formations require choosing the type of constitutive model and the model parameters based on core, log, and geophysical data during the characterization phase, with little feedback from operational observations to validate or refine these choices. It is shown that errors in the assumed constitutive response, even when informed by laboratory tests on core samples, are likely to be common, large, and underestimate the magnitude of stress change caused by injection. Recent advances in borehole-based strain instruments and borehole and surface-based tilt and displacement instruments have now enabled monitoring of the deformation of the storage system throughout its operational lifespan. This data can enable validation and refinement of the knowledge of the geomechanical properties and state of the system, but brings with it a challenge to transform the raw data into actionable knowledge. We demonstrate a method to perform a gradient-based deterministic inversion of geomechanical monitoring data. This approach allows autonomous integration of the instrument data without the need for time consuming manual interpretation and selection of updated model parameters. The approach presented is very flexible as to what type of geomechanical constitutive response can be used. The approach is easily adaptable to nonlinear physics-based constitutive models to account for common rock behaviors such as creep and plasticity. The approach also enables training of machine learning-based constitutive models by allowing back propagation of errors through the finite element calculations. This enables strongly enforcing known physics, such as conservation of momentum and continuity, while allowing data-driven models to learn the truly unknown physics such as the constitutive or petrophysical responses.

physics.geo-ph

Physics-constrained deep neural network method for estimating parameters in a redox flow battery

In this paper, we present a physics-constrained deep neural network (PCDNN) method for parameter estimation in the zero-dimensional (0D) model of the vanadium redox flow battery (VRFB). In this approach, we use deep neural networks (DNNs) to approximate the model parameters as functions of the operating conditions. This method allows the integration of the VRFB computational models as the physical constraints in the parameter learning process, leading to enhanced accuracy of parameter estimation and cell voltage prediction. Using an experimental dataset, we demonstrate that the PCDNN method can estimate model parameters for a range of operating conditions and improve the 0D model prediction of voltage compared to the 0D model prediction with constant operation-condition-independent parameters estimated with traditional inverse methods. We also demonstrate that the PCDNN approach has an improved generalization ability for estimating parameter values for operating conditions not used in the DNN training.

physics.chem-ph

An efficient epistemic uncertainty quantification algorithm for a class of stochastic models: A post-processing and domain decomposition framework

Partial differential equations (PDEs) are fundamental for theoretically describing numerous physical processes that are based on some input fields in spatial configurations. Understanding the physical process, in general, requires computational modeling of the PDE. Uncertainty in the computational model manifests through lack of precise knowledge of the input field or configuration. Uncertainty quantification (UQ) in the output physical process is typically carried out by modeling the uncertainty using a random field, governed by an appropriate covariance function. This leads to solving high-dimensional stochastic counterparts of the PDE computational models. Such UQ-PDE models require a large number of simulations of the PDE in conjunction with samples in the high-dimensional probability space, with probability distribution associated with the covariance function. Those UQ computational models having explicit knowledge of the covariance function are known as aleatoric UQ (AUQ) models. The lack of such explicit knowledge leads to epistemic UQ (EUQ) models, which typically require solution of a large number of AUQ models. In this article, using a surrogate, post-processing, and domain decomposition framework with coarse stochastic solution adaptation, we develop an offline/online algorithm for efficiently simulating a class of EUQ-PDE models.

math.NA

Highly-scalable, physics-informed GANs for learning solutions of stochastic PDEs

Uncertainty quantification for forward and inverse problems is a central challenge across physical and biomedical disciplines. We address this challenge for the problem of modeling subsurface flow at the Hanford Site by combining stochastic computational models with observational data using physics-informed GAN models. The geographic extent, spatial heterogeneity, and multiple correlation length scales of the Hanford Site require training a computationally intensive GAN model to thousands of dimensions. We develop a hierarchical scheme for exploiting domain parallelism, map discriminators and generators to multiple GPUs, and employ efficient communication schemes to ensure training stability and convergence. We developed a highly optimized implementation of this scheme that scales to 27,500 NVIDIA Volta GPUs and 4584 nodes on the Summit supercomputer with a 93.1% scaling efficiency, achieving peak and sustained half-precision rates of 1228 PF/s and 1207 PF/s.

physics.comp-ph

Engineering Structural Robustness in Power Grid Networks Susceptible to Coherent Swing Instability

Networked power grid systems are susceptible to a phenomenon known as Coherent Swing Instability (CSI), in which a subset of machines in the grid lose synchrony with the rest of the network. We develop network level evaluation metrics to (i) identify community substructures in the power grid network, (ii) determine weak points in the network that are particularly sensitive to CSI, and (iii) produce an engineering approach for the addition of transmission lines to reduce the incidences of CSI in existing networks, or design new power grid networks that are robust to CSI by their network design. For simulations on a reduced model for the American Northeast power grid, where a block of buses representing the New England region exhibit a strong propensity for CSI, we show that modifying the network's connectivity structure can markedly improve the grid's resilience to CSI. Our analysis provides a versatile diagnostic tool for evaluating the efficacy of adding lines to a power grid which is known to be prone to CSI. This is a particularly relevant problem in large-scale power systems, where improving stability and robustness to interruptions by increasing overall network connectivity is not feasible due to financial and infrastructural constraints.

physics.soc-ph

A comparative study of physics-informed neural network models for learning unknown dynamics and constitutive relations

We investigate the use of discrete and continuous versions of physics-informed neural network methods for learning unknown dynamics or constitutive relations of a dynamical system. For the case of unknown dynamics, we represent all the dynamics with a deep neural network (DNN). When the dynamics of the system are known up to the specification of constitutive relations (that can depend on the state of the system), we represent these constitutive relations with a DNN. The discrete versions combine classical multistep discretization methods for dynamical systems with neural network based machine learning methods. On the other hand, the continuous versions utilize deep neural networks to minimize the residual function for the continuous governing equations. We use the case of a fedbatch bioreactor system to study the effectiveness of these approaches and discuss conditions for their applicability. Our results indicate that the accuracy of the trained neural network models is much higher for the cases where we only have to learn a constitutive relation instead of the whole dynamics. This finding corroborates the well-known fact from scientific computing that building as much structural information is available into an algorithm can enhance its efficiency and/or accuracy.

cs.LG

Physics-Informed CoKriging: A Gaussian-Process-Regression-Based Multifidelity Method for Data-Model Convergence

In this work, we propose a new Gaussian process regression (GPR)-based multifidelity method: physics-informed CoKriging (CoPhIK). In CoKriging-based multifidelity methods, the quantities of interest are modeled as linear combinations of multiple parameterized stationary Gaussian processes (GPs), and the hyperparameters of these GPs are estimated from data via optimization. In CoPhIK, we construct a GP representing low-fidelity data using physics-informed Kriging (PhIK), and model the discrepancy between low- and high-fidelity data using a parameterized GP with hyperparameters identified via optimization. Our approach reduces the cost of optimization for inferring hyperparameters by incorporating partial physical knowledge. We prove that the physical constraints in the form of deterministic linear operators are satisfied up to an error bound. Furthermore, we combine CoPhIK with a greedy active learning algorithm for guiding the selection of additional observation locations. The efficiency and accuracy of CoPhIK are demonstrated for reconstructing the partially observed modified Branin function, reconstructing the sparsely observed state of a steady state heat transport problem, and learning a conservative tracer distribution from sparse tracer concentration measurements.

stat.ML

Physics-Information-Aided Kriging: Constructing Covariance Functions using Stochastic Simulation Models

In this work, we propose a new Gaussian process regression (GPR) method: physics information aided Kriging (PhIK). In the standard data-driven Kriging, the unknown function of interest is usually treated as a Gaussian process with assumed stationary covariance with hyperparameters estimated from data. In PhIK, we compute the mean and covariance function from realizations of available stochastic models, e.g., from realizations of governing stochastic partial differential equations solutions. Such constructed Gaussian process generally is non-stationary, and does not assume a specific form of the covariance function. Our approach avoids the optimization step in data-driven GPR methods to identify the hyperparameters. More importantly, we prove that the physical constraints in the form of a deterministic linear operator are guaranteed in the resulting prediction. We also provide an error estimate in preserving the physical constraints when errors are included in the stochastic model realizations. To reduce the computational cost of obtaining stochastic model realizations, we propose a multilevel Monte Carlo estimate of the mean and covariance functions. Further, we present an active learning algorithm that guides the selection of additional observation locations. The efficiency and accuracy of PhIK are demonstrated for reconstructing a partially known modified Branin function, studying a three-dimensional heat transfer problem and learning a conservative tracer distribution from sparse concentration measurements.

stat.ML

A diffuse-interface model for smoothed particle hydrodynamics

Diffuse-interface theory provides a foundation for the modeling and simulation of microstructure evolution in a very wide range of materials, and for the tracking/capturing of dynamic interfaces between different materials on larger scales. Smoothed particle hydrodynamics (SPH) is also widely used to simulate fluids and solids that are subjected to large deformations and have complex dynamic boundaries and/or interfaces, but no explicit interface tracking/capturing is required, even when topological changes such as fragmentation and coalescence occur, because of its Lagrangian particle nature. Here we developed an SPH model for single-component two-phase fluids that is based on diffuse-interface theory. In the model, the interface has a finite thickness and a surface tension that depend on the coefficient, k, of the gradient contribution to the Helmholtz free energy functional and the density dependent homogeneous free energy. In this model, there is no need to locate the surface (or interface) or to compute the curvature at and near the interface. One- and two-dimensional SPH simulations were used to validate the model.

physics.comp-ph

A dissipative particle dynamics model of biofilm growth

A dissipative particle dynamics (DPD) model for the quantitative simulation of biofilm growth controlled by substrate (nutrient) consumption, advective and diffusive substrate transport, and hydrodynamic interactions with fluid flow (including fragmentation and reattachment) is described. The model was used to simulate biomass growth, decay, and spreading. It predicts how the biofilm morphology depends on flow conditions, biofilm growth kinetics, the rheomechanical properties of the biofilm and adhesion to solid surfaces. The morphology of the model biofilm depends strongly on its rigidity and the magnitude of the body force that drives the fluid over the biofilm.

physics.comp-ph

Physics-informed Machine Learning Method for Forecasting and Uncertainty Quantification of Partially Observed and Unobserved States in Power Grids

We present a physics-informed Gaussian Process Regression (GPR) model to predict the phase angle, angular speed, and wind mechanical power from a limited number of measurements. In the traditional data-driven GPR method, the form of the Gaussian Process covariance matrix is assumed and its parameters are found from measurements. In the physics-informed GPR, we treat unknown variables (including wind speed and mechanical power) as a random process and compute the covariance matrix from the resulting stochastic power grid equations. We demonstrate that the physics-informed GPR method is significantly more accurate than the standard data-driven one for immediate forecasting of generators' angular velocity and phase angle. We also show that the physics-informed GPR provides accurate predictions of the unobserved wind mechanical power, phase angle, or angular velocity when measurements from only one of these variables are available. The immediate forecast of observed variables and predictions of unobserved variables can be used for effectively managing power grids (electricity market clearing, regulation actions) and early detection of abnormal behavior and faults. The physics-based GPR forecast time horizon depends on the combination of input (wind power, load, etc.) correlation time and characteristic (relaxation) time of the power grid and can be extended to short and medium-range times.

eess.SP

Sliced-Inverse-Regression-Aided Rotated Compressive Sensing Method for Uncertainty Quantification

Compressive-sensing-based uncertainty quantification methods have become a pow- erful tool for problems with limited data. In this work, we use the sliced inverse regression (SIR) method to provide an initial guess for the alternating direction method, which is used to en- hance sparsity of the Hermite polynomial expansion of stochastic quantity of interest. The sparsity improvement increases both the efficiency and accuracy of the compressive-sensing- based uncertainty quantification method. We demonstrate that the initial guess from SIR is more suitable for cases when the available data are limited (Algorithm 4). We also propose another algorithm (Algorithm 5) that performs dimension reduction first with SIR. Then it constructs a Hermite polynomial expansion of the reduced model. This method affords the ability to approximate the statistics accurately with even less available data. Both methods are non-intrusive and require no a priori information of the sparsity of the system. The effec- tiveness of these two methods (Algorithms 4 and 5) are demonstrated using problems with up to 500 random dimensions.

math.NA