SearcharxivSearch

arXiv subjects

Yanzhao Cao

Publications and source records attributed to Yanzhao Cao.

At least 19 recordsLinked to original sources

Locking analysis and high-order locking-free finite element method for three-dimensional electroporoelasticity equations

Electroporoelasticity equations couple Maxwell's equations with Biot's poroelasticity model and admit severe Poisson locking in standard conforming finite element discretizations when the Lam\'e constant is large. In this paper, we provide a rigorous analysis of the Poisson locking phenomenon for three-dimensional quasi-static electroporoelasticity equations and show that the spatial convergence order of conforming finite element approximations is reduced in the nearly incompressible regime. To eliminate this locking effect, we introduce a five-field formulation and a fully discrete high-order finite element method. We prove uniform stability with respect to large Lam\'e constant. Based on this, we establish a locking-free scheme by deriving its uniform error estimates with respect to the Lam\'e coefficient. The analysis covers the fully coupled electromagnetic-poroelastic system and applies to high-order elements in three dimensions. Extensive numerical experiments are presented to verify the theoretical convergence rates and demonstrate robustness with respect to the Lam\'e parameter.

math.NA

A Multi-Level Machine Learning Framework for Inverse Scattering Problems with Multi-Frequency Data

In this work, we propose a multi-level machine learning framework for solving inverse scattering problems with multi-frequency data. The multi-level neural network is built along the frequency axis of the scattering problem, wherein at each fixed frequency, a new level of network is added to the existing architecture to update the reconstruction. By marching through the frequency levels, the proposed multi-level computational framework is able to obtain higher-order Fourier modes of the imaging target as the depth of the neural network grows and higher-frequency data are used. Furthermore, the overall learning problem is decomposed into a sequence of simpler local tasks, each associated with a single frequency. This decomposition significantly reduces the complexity of the optimization problem and mitigates the risk of convergence to undesirable local minima, resulting in a robust and reliable training procedure for solving inverse scattering problems. We conduct various numerical experiments for the inverse source scattering problem and the inverse medium scattering problem to illustrate the effectiveness and robustness of the proposed machine learning framework. In addition, theoretical analysis in the neural tangent kernel regime shows that the proposed multi-level architecture progressively recovers the higher-order Fourier components of the imaging target.

math.NA

Strong convergence of an explicit full-discrete scheme for stochastic Burgers-Huxley equation

The strong convergence of an explicit full-discrete scheme is investigated for the stochastic Burgers-Huxley equation driven by additive space-time white noise, which possesses both Burgers-type and cubic nonlinearities. To discretize the continuous problem in space, we utilize a spectral Galerkin method. Subsequently, we introduce a nonlinear-tamed exponential integrator scheme, resulting in a fully discrete scheme. Within the framework of semigroup theory, this study provides precise estimations of the Sobolev regularity, $L^\infty$ regularity in space, and Hölder continuity in time for the mild solution, as well as for its semi-discrete and full-discrete approximations. Building upon these results, we establish moment boundedness for the numerical solution and obtain strong convergence rates in both spatial and temporal dimensions. A numerical example is presented to validate the theoretical findings.

math.NA

Maximum Principles for Partially Observed Controls of Forward SPDEs and Backward SDEs with Jumps

This work establishes two versions of the Pontryagin-type maximum principles for partially observed optimal control of coupled forward stochastic partial differential equations (FSPDEs) and backward stochastic differential equations (BSDEs) with jumps in convex control domains. The FSPDE-BSDE system is driven by cylindrical Wiener processes, finite-dimensional Brownian motions, and compensated Poisson random measures. For systems with deterministic coefficients, a direct method is employed and particular attention is focused on establishing the well-posedness of a singular backward SPDE with jumps. For systems with random coefficients, a Malliavin calculus approach is developed. The main novelty here is the establishment of the well-posedness of an operator-valued SPDE with jumps, which provides a new stochastic flow representation for linear SPDEs with jumps.

math.OC

Error estimates of a training-free diffusion model for high-dimensional sampling

Score-based diffusion models are a powerful class of generative models, but their practical use often depends on training neural networks to approximate the score function. Training-free diffusion models provide an attractive alternative by exploiting analytically tractable score functions, and have recently enabled supervised learning of efficient end-to-end generative samplers. Despite their empirical success, the training-free diffusion models lack rigorous and numerically verifiable error estimates. In this work, we develop a comprehensive error analysis for a class of training-free diffusion models used to generate labeled data for supervised learning of generative samplers. By exploiting the availability of the exact score function for Gaussian mixture models, our analysis avoids propagating score-function approximation errors through the reverse-time diffusion process and recovers classical convergence rates for ODE discretization schemes, such as first-order convergence for the Euler method. Moreover, the resulting error bounds exhibit favorable dimension dependence, scaling as $O(d)$ in the $\ell_2$ norm and $O(\log d)$ in the $\ell_\infty$ norm. Importantly, the proposed error estimates are fully numerically verifiable with respect to both time-step size and dimensionality, thereby bridging the gap between theoretical analysis and observed numerical behavior.

math.NA

Learning Stochastic Hamiltonian Systems via Stochastic Generating Function Neural Network

In this paper we propose a novel neural network model for learning stochastic Hamiltonian systems (SHSs) from observational data, termed the stochastic generating function neural network (SGFNN). SGFNN preserves symplectic structure of the underlying stochastic Hamiltonian system and produces symplectic predictions. Our model utilizes the autoencoder framework to identify the randomness of the latent system by the encoder network, and detects the stochastic generating function of the system through the decoder network based on the random variables extracted from the encoder. Symplectic predictions can then be generated by the stochastic generating function. Numerical experiments are performed on several stochastic Hamiltonian systems, varying from additive to multiplicative, and from separable to non-separable SHSs with single or multiple noises. Compared with the benchmark stochastic flow map learning (sFML) neural network, our SGFNN model exhibits higher accuracy across various prediction metrics, especially in long-term predictions, with the property of maintaining the symplectic structure of the underlying SHSs.

math.DS

Generative AI Models for Learning Flow Maps of Stochastic Dynamical Systems in Bounded Domains

Simulating stochastic differential equations (SDEs) in bounded domains, presents significant computational challenges due to particle exit phenomena, which requires accurate modeling of interior stochastic dynamics and boundary interactions. Despite the success of machine learning-based methods in learning SDEs, existing learning methods are not applicable to SDEs in bounded domains because they cannot accurately capture the particle exit dynamics. We present a unified hybrid data-driven approach that combines a conditional diffusion model with an exit prediction neural network to capture both interior stochastic dynamics and boundary exit phenomena. Our ML model consists of two major components: a neural network that learns exit probabilities using binary cross-entropy loss with rigorous convergence guarantees, and a training-free diffusion model that generates state transitions for non-exiting particles using closed-form score functions. The two components are integrated through a probabilistic sampling algorithm that determines particle exit at each time step and generates appropriate state transitions. The performance of the proposed approach is demonstrated via three test cases: a one-dimensional simplified problem for theoretical verification, a two-dimensional advection-diffusion problem in a bounded domain, and a three-dimensional problem of interest to magnetically confined fusion plasmas.

stat.ML

Optimal Control of Stochastic Partial Differential Equations with Partial Observations: Stochastic Maximum Principles and Numerical Approximation

This work establishes a general stochastic maximum principle for partially observed optimal control of semi-linear stochastic partial differential equations in a nonconvex control domain. The state evolves in a Hilbert space driven by a cylindrical Wiener process and finitely many Brownian motions, while observations are in an Euclidean space having correlated noise. For convex control domain and diffusion coefficients in the state being control-independent, numerical algorithms are developed to solve the partially observed optimal control problems using stochastic gradient descent algorithm combined with finite element approximations and the branching filtering algorithm. Numerical experiments are conducted for demonstration.

math.OC

Numerical approximations for partially observed optimal control of stochastic partial differential equations

In this paper, we study numerical approximations for optimal control of a class of stochastic partial differential equations with partial observations. The system state evolves in a Hilbert space, whereas observations are given in finite-dimensional space $\rr^d$. We begin by establishing stochastic maximum principles (SMPs) for such problems, where the system state is driven by a cylindrical Wiener process. The corresponding adjoint equations are characterized by backward stochastic partial differential equations. We then develop numerical algorithms to solve the partially observed optimal control. Our approach combines the stochastic gradient descent method, guided by the SMP, with a particle filtering algorithm to estimate the conditional distributions of the state of the system. Finally, we demonstrate the effectiveness of our proposed algorithm through numerical experiments.

math.OC

Weak Convergence Analysis for the Finite Element Approximation to Stochastic Allen-Cahn Equation Driven by Multiplicative White Noise

In this paper, we aim to study the optimal weak convergence order for the finite element approximation to a stochastic Allen-Cahn equation driven by multiplicative white noise. We first construct an auxiliary equation based on the splitting-up technique and derive prior estimates for the corresponding Kolmogorov equation and obtain the strong convergence order of 1 in time between the auxiliary and exact solutions. Then, we prove the optimal weak convergence order of the finite element approximation to the stochastic Allen-Cahn equation by deriving the weak convergence order between the finite element approximation and the auxiliary solution via the theory of Kolmogorov equation and Malliavin calculus. Finally, we present a numerical experiment to illustrate the theoretical analysis.

math.NA

Learning Hamiltonian Systems with Pseudo-symplectic Neural Network

In this paper, we introduces a Pseudo-Symplectic Neural Network (PSNN) for learning general Hamiltonian systems (both separable and non-separable) from data. To address the limitations of existing structure-preserving methods (e.g., implicit symplectic integrators restricted to separable systems or explicit approximations requiring high computational costs), PSNN integrates an explicit pseudo-symplectic integrator as its dynamical core, achieving nearly exact symplecticity with minimal structural error. Additionally, the authors propose learnable Padé-type activation functions based on Padé approximation theory, which empirically outperform classical ReLU, Taylor-based activations, and PAU. By combining these innovations, PSNN demonstrates superior performance in learning and forecasting diverse Hamiltonian systems (e.g., modified pendulum, galactic dynamics), surpassing state-of-the-art models in accuracy, long-term stability, and energy preservation, while requiring shorter training time, fewer samples, and reduced parameters. This framework bridges the gap between computational efficiency and geometric structure preservation in Hamiltonian system modeling.

math.NA

Splitting finite element approximations for quasi-static electroporoelasticity equations

The electroporoelasticity model, which couples Maxwell's equations with Biot's equations, plays a critical role in applications such as water conservancy exploration, earthquake early warning, and various other fields. This work focuses on investigating its well-posedness and analyzing error estimates for a splitting backward Euler finite element method. We first define a weak solution consistent with the finite element framework. Then, we prove the uniqueness and existence of such a solution using the Galerkin method and derive a priori estimates for high-order regularity. Using a splitting technique, we define an approximate splitting solution and analyze its convergence order. Next, we apply Nedelec's curl-conforming finite elements, Lagrange elements, and the backward Euler method to construct a fully discretized scheme. We demonstrate the stability of the splitting numerical solution and provide error estimates for its convergence order in both temporal and spatial variables. Finally, we present numerical experiments to validate the theoretical results, showing that our method significantly reduces computational complexity compared to the classical finite element method.

math.NA

Conditional Pseudo-Reversible Normalizing Flow for Surrogate Modeling in Quantifying Uncertainty Propagation

We introduce a conditional pseudo-reversible normalizing flow for constructing surrogate models of a physical model polluted by additive noise to efficiently quantify forward and inverse uncertainty propagation. Existing surrogate modeling approaches usually focus on approximating the deterministic component of physical model. However, this strategy necessitates knowledge of noise and resorts to auxiliary sampling methods for quantifying inverse uncertainty propagation. In this work, we develop the conditional pseudo-reversible normalizing flow model to directly learn and efficiently generate samples from the conditional probability density functions. The training process utilizes dataset consisting of input-output pairs without requiring prior knowledge about the noise and the function. Our model, once trained, can generate samples from any conditional probability density functions whose high probability regions are covered by the training set. Moreover, the pseudo-reversibility feature allows for the use of fully-connected neural network architectures, which simplifies the implementation and enables theoretical analysis. We provide a rigorous convergence analysis of the conditional pseudo-reversible normalizing flow model, showing its ability to converge to the target conditional probability density function using the Kullback-Leibler divergence. To demonstrate the effectiveness of our method, we apply it to several benchmark tests and a real-world geologic carbon storage problem.

cs.LG

Diffusion-Model-Assisted Supervised Learning of Generative Models for Density Estimation

We present a supervised learning framework of training generative models for density estimation. Generative models, including generative adversarial networks, normalizing flows, variational auto-encoders, are usually considered as unsupervised learning models, because labeled data are usually unavailable for training. Despite the success of the generative models, there are several issues with the unsupervised training, e.g., requirement of reversible architectures, vanishing gradients, and training instability. To enable supervised learning in generative models, we utilize the score-based diffusion model to generate labeled data. Unlike existing diffusion models that train neural networks to learn the score function, we develop a training-free score estimation method. This approach uses mini-batch-based Monte Carlo estimators to directly approximate the score function at any spatial-temporal location in solving an ordinary differential equation (ODE), corresponding to the reverse-time stochastic differential equation (SDE). This approach can offer both high accuracy and substantial time savings in neural network training. Once the labeled data are generated, we can train a simple fully connected neural network to learn the generative model in the supervised manner. Compared with existing normalizing flow models, our method does not require to use reversible neural networks and avoids the computation of the Jacobian matrix. Compared with existing diffusion models, our method does not need to solve the reverse-time SDE to generate new samples. As a result, the sampling efficiency is significantly improved. We demonstrate the performance of our method by applying it to a set of 2D datasets as well as real data from the UCI repository.

cs.LG

A pseudo-reversible normalizing flow for stochastic dynamical systems with various initial distributions

We present a pseudo-reversible normalizing flow method for efficiently generating samples of the state of a stochastic differential equation (SDE) with different initial distributions. The primary objective is to construct an accurate and efficient sampler that can be used as a surrogate model for computationally expensive numerical integration of SDE, such as those employed in particle simulation. After training, the normalizing flow model can directly generate samples of the SDE's final state without simulating trajectories. Existing normalizing flows for SDEs depend on the initial distribution, meaning the model needs to be re-trained when the initial distribution changes. The main novelty of our normalizing flow model is that it can learn the conditional distribution of the state, i.e., the distribution of the final state conditional on any initial state, such that the model only needs to be trained once and the trained model can be used to handle various initial distributions. This feature can provide a significant computational saving in studies of how the final state varies with the initial distribution. We provide a rigorous convergence analysis of the pseudo-reversible normalizing flow model to the target probability density function in the Kullback-Leibler divergence metric. Numerical experiments are provided to demonstrate the effectiveness of the proposed normalizing flow model.

math.NA

Convergence Analysis for Training Stochastic Neural Networks via Stochastic Gradient Descent

In this paper, we carry out numerical analysis to prove convergence of a novel sample-wise back-propagation method for training a class of stochastic neural networks (SNNs). The structure of the SNN is formulated as discretization of a stochastic differential equation (SDE). A stochastic optimal control framework is introduced to model the training procedure, and a sample-wise approximation scheme for the adjoint backward SDE is applied to improve the efficiency of the stochastic optimal control solver, which is equivalent to the back-propagation for training the SNN. The convergence analysis is derived with and without convexity assumption for optimization of the SNN parameters. Especially, our analysis indicates that the number of SNN training steps should be proportional to the square of the number of layers in the convex optimization case. Numerical experiments are carried out to validate the analysis results, and the performance of the sample-wise back-propagation method for training SNNs is examined by benchmark machine learning examples.

math.NA

Numerical analysis of a time discretized method for nonlinear filtering problem with Lévy process observations

In this paper, we consider a nonlinear filtering model with observations driven by correlated Wiener processes and point processes. We first derive a Zakai equation whose solution is a unnormalized probability density function of the filter solution. Then we apply a splitting-up technique to decompose the Zakai equation into three stochastic differential equations, based on which we construct a splitting-up approximate solution and prove its half-order convergence. Furthermore, we apply a finite difference method to construct a time semi-discrete approximate solution to the splitting-up system and prove its half-order convergence to the exact solution of the Zakai equation. Finally, we present some numerical experiments to demonstrate the theoretical analysis.

math.NA

A probabilistic scheme for semilinear nonlocal diffusion equations with volume constraints

This work presents a probabilistic scheme for solving semilinear nonlocal diffusion equations with volume constraints and integrable kernels. The nonlocal model of interest is defined by a time-dependent semilinear partial integro-differential equation (PIDE), in which the integro-differential operator consists of both local convection-diffusion and nonlocal diffusion operators. Our numerical scheme is based on the direct approximation of the nonlinear Feynman-Kac formula that establishes a link between nonlinear PIDEs and stochastic differential equations. The exploitation of the Feynman-Kac representation successfully avoids solving dense linear systems arising from nonlocality operators. Compared with existing stochastic approaches, our method can achieve first-order convergence after balancing the temporal and spatial discretization errors, which is a significant improvement of existing probabilistic/stochastic methods for nonlocal diffusion problems. Error analysis of our numerical scheme is established. The effectiveness of our approach is shown in two numerical examples. The first example considers a three-dimensional nonlocal diffusion equation to numerically verify the error analysis results. The second example presents a physics problem motivated by the study of heat transport in magnetically confined fusion plasmas.

math.NA