SearcharxivSearch

arXiv subjects

Tan Bui-Thanh

Publications and source records attributed to Tan Bui-Thanh.

At least 19 recordsLinked to original sources

Second-order consistency for learning chaotic dynamics via randomized Jacobian matching

Short-horizon accuracy does not ensure that a learned chaotic system has correct long-time dynamics. Trajectory (zeroth-order) matching constrains vector-field values, and Jacobian (first-order) matching constrains local tangent dynamics, but neither determines how the Jacobian varies away from supervised states, so a model can be locally accurate while drifting toward spurious attractors and distorting long-time statistics. We show that second-order supervision mitigates these failures. Because forming full Hessian tensors is computationally prohibitive in high dimensions, we propose model-constrained randomized Jacobian matching, which compares the Jacobians of the true and learned vector fields at randomly perturbed inputs. A Taylor expansion shows that the expected randomized Jacobian loss decomposes into the Jacobian mismatch plus a Hessian mismatch scaled by the noise variance, implicitly enforcing second-order consistency at $O(d^2)$ memory cost without forming the $O(d^3)$ Hessian tensor. In Lorenz 63 with minimal temporal supervision, second-order supervision reduces invariant-measure error and Lyapunov-spectrum MSE, and recovers the constant Hessian norm of the true bilinear field. Across five training seeds, explicit Hessian matching produces catastrophic Lyapunov outliers for four of five seeds, whereas randomized Jacobian matching produces none among 5,000 on-attractor rollouts, and it attains the largest threshold in a directional capture scan. In coupled Lorenz 96, first-order methods enter spurious high-amplitude regimes as forcing increases, while second-order methods retain accurate marginals. Randomized Jacobian matching costs about the same as explicit Hessian matching on Lorenz 63 and 40% less on Lorenz 96, with no reference Hessian evaluations during training.

math.NA

Dimension Bridging for 3D RANS with Neural Network Accelerated Gaussian Functional Regression

In many computational science and engineering problems, repeatedly solving fully resolved physics-based models to design for a quantity of interest (QoI) can quickly become intractable, requiring the use of low-fidelity models to predict the same QoI but introduce errors where some features are neglected or are otherwise inaccurately resolved. We use Gaussian Functional Regression (GFR) to learn a correction to a 2D Reynolds-Averaged Navier-Stokes (RANS) model to predict the aerodynamic coefficients from a 3D RANS model. This model pair has a disparity in the governing physics from the reduced dimensionality, a previously unexplored application for GFR. Empirically, our results show that with a proper choice of low-dimensional (LD) model, the proposed kernel allows for the use of fewer high-dimensional (HD) evaluations to regress a response surface to the same level of accuracy as standard stationary kernels. Moreover, the new kernel provides more informative uncertainty quantification, which we show is advantageous when used to drive an adaptive sampling algorithm. Finally, we propose a novel neural network accelerated kernel, which we show offers predictions in good agreement while speeding up evaluations by millions of times in wall clock measurements, bringing the computational budget within the real-time regime.

cs.CE

An optimal control approach for neural network architecture adaptation with a posteriori error estimation

This work presents a novel approach for adapting neural network architecture along the depth based on a posteriori error estimation. By formulating neural network training as a continuous-time optimal control problem, we derive rigorous error estimates that quantify how approximation error distributes across network layers. This error decomposition enables a principled depth adaptation strategy: new layers are inserted at locations of maximum estimated error, allowing the network to efficiently capture complex, nonlinear variations in the underlying problem. Our framework introduces a novel network architecture that treats weights and biases as piecewise linear functions varying across layers, with the error estimator bounding the discrepancy between this discrete representation and the true continuous optimal control solution. The approach leverages dual weighted residual methodology from finite element analysis to derive computable upper bounds on the functional error. A key theoretical contribution is the derivation of explicit error bounds that decompose the total approximation error into interval-wise contributions, providing a rigorous basis for targeted architecture refinement. We demonstrate the effectiveness of our method on scientific datasets, including learning the observable-to-parameter map for the Navier-Stokes equation. Numerical results reveal that our approach consistently outperforms existing architecture adaptation methods in terms of generalization performance.

cs.LG

Laplace-Fisher Gate Identities for Optimal Matrix-Gated Blended Score Estimation

Sampling from an unnormalized target density by reversing an Ornstein-Uhlenbeck diffusion requires the score of each noise-perturbed marginal law. Two exact identities are available: Tweedie's identity and a target-score identity, each yielding unbiased finite-reference score estimators for the OU-marginal score. Score estimators induced by scalar blends of Tweedie and TSI score estimators can reduce variance, but they are too rigid for singular or strongly anisotropic targets. We formulate blended score estimation as a conditional risk-minimization problem over matrix valued blending coefficients, referred to as gates. Our central result is to show the optimal matrix valued gate for blended score estimation is given \[ G_\star(y,t) = α_t^2 \left(α_t^2 I_d + γ_t\, \mathbb{E}[H_0(X_0)\mid Y_t=y] \right)^{-1}, \qquad H_0=-\nabla^2\log p_0 .\] Here $α_t = e^{-t}$ and $γ_t = 1-e^{-2t}$ are the OU coefficients, and the conditional expectation is under the OU posterior of $X_0$ given $Y_t=y$. We call this formula the \emph{Laplace-Fisher Gate Identity} (\LFGI{}). Because the Tweedie-TSI disagreement has conditional mean zero, the gate changes the score-estimator variance but not its expected value. We derive the variance-optimal matrix gate, record the Gaussian special case, and establish finite-reference consistency and stability bounds for estimating the gate from weighted reference samples. We then use the finite-reference LFGI score estimator for normalized density evaluation in Bayesian inverse problems. In regimes where MCMC pilot samples and derivative information are already available, LFGI uses those byproducts to construct a normalized surrogate for the posterior density. The resulting surrogate supplies information that the MCMC samples alone do not provide: posterior-energy evaluation, model-evidence estimation, and downstream density-based diagnostics. On a PDE-constrained inverse-problem benchmark, the LFGI surrogate improves posterior-density calibration and sampling diagnostics relative to the other tested score-estimator classes. Experiments using LFGI with known model evidence check absolute evidence calibration in both Gaussian and non-Gaussian settings.

math.ST

The AI Research Assistant: Promise, Peril, and a Proof of Concept

Can artificial intelligence truly contribute to creative mathematical research, or does it merely automate routine calculations while introducing risks of error? We provide empirical evidence through a detailed case study: the discovery of novel error representations and bounds for Hermite quadrature rules via systematic human-AI collaboration. Working with multiple AI assistants, we extended results beyond what manual work achieved, formulating and proving several theorems with AI assistance. The collaboration revealed both remarkable capabilities and critical limitations. AI excelled at algebraic manipulation, systematic proof exploration, literature synthesis, and LaTeX preparation. However, every step required rigorous human verification, mathematical intuition for problem formulation, and strategic direction. We document the complete research workflow with unusual transparency, revealing patterns in successful human-AI mathematical collaboration and identifying failure modes researchers must anticipate. Our experience suggests that, when used with appropriate skepticism and verification protocols, AI tools can meaningfully accelerate mathematical discovery while demanding careful human oversight and deep domain expertise.

cs.AI

Rendezvous Planning from Sparse Observations of Optimally Controlled Targets

We develop a probabilistic framework for \emph{rendezvous planning}: given sparse, noisy observations of a fast-moving target, plan rendezvous spatiotemporal coordinates for a set of significantly slower seeking agents. The unknown target trajectory is estimated under uncertain dynamics using a filtering approach that combines a kernel-based maximum a posteriori estimation with Gaussian process correction, producing a mixture over trajectory hypotheses. This estimate is used to select spatiotemporal rendezvous points that maximize the probability of successful rendezvous. Points are chosen sequentially by greedily minimizing failure probability in the current belief space, which is updated after each step by conditioning on unsuccessful rendezvous attempts. We show that the failure-conditioned update correctly captures the posterior belief for subsequent decisions, ensuring that each step in the greedy sequence is informed by a statistically consistent representation of the remaining search space, and derive the corresponding Bayesian updates incorporating temporal correlations intrinsic to the trajectory model. This result provides a systematic framework for planning under uncertainty in applications of autonomous rendezvous such as unmanned aerial vehicle refueling, spacecraft servicing, autonomous surface vessel operations, search and rescue missions, and missile defense. In each, the motion of the target entity can be modeled using a system of differential equations undergoing optimal control for a chosen objective, in our example case Hamilton--Jacobi--Bellman solutions for minimum arrival time of a Dubins car with uncertain turning radius and destination.

math.OC

Generalization Limits of In-Context Operator Networks for Higher-Order Partial Differential Equations

We investigate the generalization capabilities of In-Context Operator Networks (ICONs), a new class of operator networks that build on the principles of in-context learning, for higher-order partial differential equations. We extend previous work by expanding the type and scope of differential equations handled by the foundation model. We demonstrate that while processing complex inputs requires some new computational methods, the underlying machine learning techniques are largely consistent with simpler cases. Our implementation shows that although point-wise accuracy degrades for higher-order problems like the heat equation, the model retains qualitative accuracy in capturing solution dynamics and overall behavior. This demonstrates the model's ability to extrapolate fundamental solution characteristics to problems outside its training regime.

cs.LG

Topological derivative approach for deep neural network architecture adaptation

This work presents a novel algorithm for progressively adapting neural network architecture along the depth. In particular, we attempt to address the following questions in a mathematically principled way: i) Where to add a new capacity (layer) during the training process? ii) How to initialize the new capacity? At the heart of our approach are two key ingredients: i) the introduction of a ``shape functional" to be minimized, which depends on neural network topology, and ii) the introduction of a topological derivative of the shape functional with respect to the neural network topology. Using an optimal control viewpoint, we show that the network topological derivative exists under certain conditions, and its closed-form expression is derived. In particular, we explore, for the first time, the connection between the topological derivative from a topology optimization framework with the Hamiltonian from optimal control theory. Further, we show that the optimality condition for the shape functional leads to an eigenvalue problem for deep neural architecture adaptation. Our approach thus determines the most sensitive location along the depth where a new layer needs to be inserted during the training phase and the associated parametric initialization for the newly added layer. We also demonstrate that our layer insertion strategy can be derived from an optimal transport viewpoint as a solution to maximizing a topological derivative in $p$-Wasserstein space, where $p>= 1$. Numerical investigations with fully connected network, convolutional neural network, and vision transformer on various regression and classification problems demonstrate that our proposed approach can outperform an ad-hoc baseline network and other architecture adaptation strategies. Further, we also demonstrate other applications of topological derivative in fields such as transfer learning.

cs.LG

From an Elementary Proof of Error Representation for Hermite Quadrature to a Rediscovery of Legendre Polynomials and Rodrigues Formula

We generalize two-point interpolatory Hermite quadrature to functions with available values and the first (n-1) derivatives at both end points. Armed with integration by parts in the reverse form we provide an elementary derivation of an exact error represenation of Hermite quadrature rule. This approach possesses several advantages over the classical approaches: i) Only integration by parts is needed for the derivation; ii) the error representation requires much milder regularity, namely the existence of nth-order derivative rather than a (2n)th-order derivative of the function under consideration. As a result, our error formula is valid for less regular functions for which the classical ones are not valid; iii) our approach rediscovers Legendre polynomials and more interestingly it provides a surprisingly elegant relation between Legendre polynomial and Hermite interpolation. In particular, Legendre polynomials are precisely the error kernels for interpolatory Hermite quadrature rules; and iv) We also rediscover the Rodrigues formula for Legendre polynomials as part of our findings. For those who are interested in a different proof of the exact error representation for Hermite quadrature rule, we provide an alternative proof using the Peano kernel theorem. We also provide a composite interpolatory Hermite quadrature rule for practical applications.

math.NA

A New Look at the Ensemble Kalman Filter for Inverse Problems: Duality, Non-Asymptotic Analysis and Convergence Acceleration

This work presents new results and understanding of the Ensemble Kalman filter (EnKF) for inverse problems. In particular, using a Lagrangian dual perspective we show that EnKF can be derived from the sample average approximation (SAA) of the Lagrangian dual function. The beauty of this new duality perspective is that it facilitates us to prove and numerically verify a novel non-asymptotic convergence result for the EnKF. Motivated by the new perspective, we also present a new convergence improvement strategy for the Ensemble Kalman Inversion Algorithm (EnKI), which is an iterative version of the EnKF for inverse problems. In particular, we propose an adaptive multiplicative correction to the sample covariance matrix at each iteration and we call this new algorithm as EnKI-MC (I). Based on the new duality perspective, we derive an expression for the optimal correction factor at each iteration of the EnKI algorithm to accelerate the convergence. In addition, we also consider an ensemble specific multiplicative covariance correction strategy (EnKI-MC (II)) where a different correction is employed for each ensemble. By viewing EnKI through the lens of fixed-point iteration, we also provide theoretical results that guarantees the convergence of EnKI-MC (I) algorithm. Numerical investigations for the deconvolution problem, initial condition inversion in advection-convection problem, initial condition inversion in a Lorenz 96 model, and inverse problem constrained by elliptic partial differential equation are conducted to verify the non-asymptotic results for EnKF and to assess the performance of convergence improvement strategies for EnKI. The numerical results suggest that the proposed strategies for EnKI not only led to faster convergence in comparison to the currently employed techniques but also better quality solutions at termination of the algorithm.

math.NA

Variance-Reduced Diffusion Sampling via Target Score Identity

We study variance reduction for score estimation and diffusion-based sampling in settings where the clean (target) score is available or can be approximated. Starting from the Target Score Identity (TSI), which expresses the noisy marginal score as a conditional expectation of the target score under the forward diffusion, we develop: (i) a plug-and-play nonparametric self-normalized importance sampling estimator compatible with standard reverse-time solvers, (ii) a variance-minimizing \emph{state- and time-dependent} blending rule between Tweedie-type and TSI estimators together with an anti-correlation analysis, (iii) a data-only extension based on locally fitted proxy scores, and (iv) a likelihood-tilting extension to Bayesian inverse problems. We also propose a \emph{Critic--Gate} distillation scheme that amortizes the state-dependent blending coefficient into a neural gate. Experiments on synthetic targets and PDE-governed inverse problems demonstrate improved sample quality for a fixed simulation budget.

stat.ML

LiLaN: A Linear Latent Network as the Solution Operator for Real-Time Solutions to Stiff Nonlinear Ordinary Differential Equations

Solving stiff ordinary differential equations (StODEs) requires sophisticated numerical solvers, which are often computationally expensive. In general, traditional explicit time integration schemes with restricted time step sizes are not suitable for StODEs, and one must resort to costly implicit methods. On the other hand, state-of-the-art machine learning based methods, such as Neural ODE, poorly handle the timescale separation of various elements of the solutions to StODEs, while still requiring expensive implicit/explicit integration at inference time. In this work, we propose a linear latent network (LiLaN) approach in which the dynamics in the latent space can be integrated analytically, and thus numerical integration is completely avoided. At the heart of LiLaN are the following key ideas: i) two encoder networks to encode the initial condition together with parameters of the ODE to the slope and the initial condition for the latent dynamics, respectively. Since the latent dynamics, by design, are linear, the solution can be evaluated analytically; ii) a neural network to map the physical time to latent times, one for each latent variable. Finally, iii) a decoder network to decode the latent solution to the physical solution at the corresponding physical time. We provide a universal approximation theorem for the proposed LiLaN approach, showing that it can approximate the solution of any stiff nonlinear system on a compact set to any degree of accuracy epsilon. We also show an interesting fact that the dimension of the latent dynamical system in LiLaN is independent of epsilon. Numerical results on the "Robertson Stiff Chemical Kinetics Model," "Plasma Collisional-Radiative Model," and "Allen-Cahn" and "Cahn-Hilliard" PDEs suggest that LiLaN outperformed state-of-the-art machine learning approaches for handling stiff ordinary and partial differential equations.

stat.ML

Improvements on uncertainty quantification with variational autoencoders

Inverse problems aim to determine model parameters of a mathematical problem from given observational data. Neural networks can provide an efficient tool to solve these problems. In the context of Bayesian inverse problems, Uncertainty Quantification Variational AutoEncoders (UQ-VAE), a class of neural networks, approximate the posterior distribution mean and covariance of model parameters. This allows for both the estimation of the parameters and their uncertainty in relation to the observational data. In this work, we propose a novel loss function for training UQ-VAEs, which includes, among other modifications, the removal of a sample mean term from an already existing one. This modification improves the accuracy of UQ-VAEs, as the original theoretical result relies on the convergence of the sample mean to the expected value (a condition that, in high dimensional parameter spaces, requires a prohibitively large number of samples due to the curse of dimensionality). Avoiding the computation of the sample mean significantly reduces the training time in high dimensional parameter spaces compared to previous literature results. Under this new formulation, we establish a new theoretical result for the approximation of the posterior mean and covariance for general mathematical problems. We validate the effectiveness of UQ-VAEs through three benchmark numerical tests: a Poisson inverse problem, a non affine inverse problem and a 0D cardiocirculatory model, under the two clinical scenarios of systemic hypertension and ventricular septal defect. For the latter case, we perform forward uncertainty quantification.

math.NA

TAEN: A Model-Constrained Tikhonov Autoencoder Network for Forward and Inverse Problems

Efficient real-time solvers for forward and inverse problems are essential in engineering and science applications. Machine learning surrogate models have emerged as promising alternatives to traditional methods, offering substantially reduced computational time. Nevertheless, these models typically demand extensive training datasets to achieve robust generalization across diverse scenarios. While physics-based approaches can partially mitigate this data dependency and ensure physics-interpretable solutions, addressing scarce data regimes remains a challenge. Both purely data-driven and physics-based machine learning approaches demonstrate severe overfitting issues when trained with insufficient data. We propose a novel Tikhonov autoencoder model-constrained framework, called TAE, capable of learning both forward and inverse surrogate models using a single arbitrary observation sample. We develop comprehensive theoretical foundations including forward and inverse inference error bounds for the proposed approach for linear cases. For comparative analysis, we derive equivalent formulations for pure data-driven and model-constrained approach counterparts. At the heart of our approach is a data randomization strategy, which functions as a generative mechanism for exploring the training data space, enabling effective training of both forward and inverse surrogate models from a single observation, while regularizing the learning process. We validate our approach through extensive numerical experiments on two challenging inverse problems: 2D heat conductivity inversion and initial condition reconstruction for time-dependent 2D Navier-Stokes equations. Results demonstrate that TAE achieves accuracy comparable to traditional Tikhonov solvers and numerical forward solvers for both inverse and forward problems, respectively, while delivering orders of magnitude computational speedups.

cs.LG

A Model-Constrained Discontinuous Galerkin Network (DGNet) for Compressible Euler Equations with Out-of-Distribution Generalization

Real-time accurate solutions of large-scale complex dynamical systems are critically needed for control, optimization, uncertainty quantification, and decision-making in practical engineering and science applications, particularly in digital twin contexts. In this work, we develop a model-constrained discontinuous Galerkin Network (DGNet) approach, a significant extension to our previous work [Model-constrained Tagent Slope Learning Approach for Dynamical Systems], for compressible Euler equations with out-of-distribution generalization. The core of DGNet is the synergy of several key strategies: (i) leveraging time integration schemes to capture temporal correlation and taking advantage of neural network speed for computation time reduction; (ii) employing a model-constrained approach to ensure the learned tangent slope satisfies governing equations; (iii) utilizing a GNN-inspired architecture where edges represent Riemann solver surrogate models and nodes represent volume integration correction surrogate models, enabling capturing discontinuity capability, aliasing error reduction, and mesh discretization generalizability; (iv) implementing the input normalization technique that allows surrogate models to generalize across different initial conditions, geometries, meshes, boundary conditions, and solution orders; and (v) incorporating a data randomization technique that not only implicitly promotes agreement between surrogate models and true numerical models up to second-order derivatives, ensuring long-term stability and prediction capacity, but also serves as a data generation engine during training, leading to enhanced generalization on unseen data. To validate the effectiveness, stability, and generalizability of our novel DGNet approach, we present comprehensive numerical results for 1D and 2D compressible Euler equation problems.

stat.ML

An Adaptive and Stability-Promoting Layerwise Training Approach for Sparse Deep Neural Network Architecture

This work presents a two-stage adaptive framework for progressively developing deep neural network (DNN) architectures that generalize well for a given training data set. In the first stage, a layerwise training approach is adopted where a new layer is added each time and trained independently by freezing parameters in the previous layers. We impose desirable structures on the DNN by employing manifold regularization, sparsity regularization, and physics-informed terms. We introduce a epsilon-delta stability-promoting concept as a desirable property for a learning algorithm and show that employing manifold regularization yields a epsilon-delta stability-promoting algorithm. Further, we also derive the necessary conditions for the trainability of a newly added layer and investigate the training saturation problem. In the second stage of the algorithm (post-processing), a sequence of shallow networks is employed to extract information from the residual produced in the first stage, thereby improving the prediction accuracy. Numerical investigations on prototype regression and classification problems demonstrate that the proposed approach can outperform fully connected DNNs of the same size. Moreover, by equipping the physics-informed neural network (PINN) with the proposed adaptive architecture strategy to solve partial differential equations, we numerically show that adaptive PINNs not only are superior to standard PINNs but also produce interpretable hidden layers with provable stability. We also apply our architecture design strategy to solve inverse problems governed by elliptic partial differential equations.

cs.LG

A Divergence-Free and $H(div)$-Conforming Embedded-Hybridized DG Method for the Incompressible Resistive MHD equations

We present a divergence-free and $H(div)$-conforming hybridized discontinuous Galerkin (HDG) method and a computationally efficient variant called embedded-HDG (E-HDG) for solving stationary incompressible viso-resistive magnetohydrodynamic (MHD) equations. The proposed E-HDG approach uses continuous facet unknowns for the vector-valued solutions (velocity and magnetic fields) while it uses discontinuous facet unknowns for the scalar variable (pressure and magnetic pressure). This choice of function spaces makes E-HDG computationally far more advantageous, due to the much smaller number of degrees of freedom, compared to the HDG counterpart. The benefit is even more significant for three-dimensional/high-order/fine mesh scenarios. On simplicial meshes, the proposed methods with a specific choice of approximation spaces are well-posed for linear(ized) MHD equations. For nonlinear MHD problems, we present a simple approach exploiting the proposed linear discretizations by using a Picard iteration. The beauty of this approach is that the divergence-free and $H(div)$-conforming properties of the velocity and magnetic fields are automatically carried over for nonlinear MHD equations. We study the accuracy and convergence of our E-HDG method for both linear and nonlinear MHD cases through various numerical experiments, including two- and three-dimensional problems with smooth and singular solutions. The numerical examples show that the proposed methods are pressure robust, and the divergence of the resulting velocity and magnetic fields is machine zero for both smooth and singular problems.

math.NA

Use of mobile phone sensing data to estimate residence and mobility times in urban patches during the COVID-19 epidemic: The case of the 2020 outbreak in Hermosillo, Mexico

It is often necessary to introduce the main characteristics of population mobility dynamics to model critical social phenomena such as the economy, violence, transmission of information, or infectious diseases. In this work, we focus on modeling and inferring urban population mobility using the geospatial data of its inhabitants. The objective is to estimate mobility and times inhabitants spend in the areas of interest, such as zip codes and census geographical areas. The proposed method uses the Brownian bridge model for animal movement in ecology. We illustrate its possible applications using mobile phone GPS data in 2020 from the city of Hermosillo, Sonora, in Mexico. We incorporate the estimated residence-mobility matrix into a multi-patch compartmental SEIR model to assess the effect of mobility changes due to governmental interventions

stat.AP