SearcharxivSearch

arXiv subjects

John Harlim

Publications and source records attributed to John Harlim.

At least 19 recordsLinked to original sources

Diffusion Maps Kernel Ridge Regression

In this paper, we study kernel ridge regression using the data-driven diffusion maps (DM) kernel, which is constructed through algebraic manipulations of the diffusion maps algorithm. Under the assumptions that the data lie on a manifold and are sampled uniformly, we prove that the appropriately scaled DM kernel converges uniformly to the heat kernel on the manifold for sufficiently large times as the dataset size increases and the kernel bandwidth is scaled appropriately. Consequently, the limiting Reproducing Kernel Hilbert Space (RKHS) induced by the DM kernel coincides with the RKHS associated with the heat kernel on the manifold. We further show that the RKHS induced by the DM kernel is isometrically isomorphic to the RKHS of the Gaussian kernel, which is continuously embedded in the RKHS of a Mat\'ern kernel whose norm is equivalent to an appropriate Sobolev norm. This result implies that standard risk bounds for kernel ridge regression applicable to Mat\'ern kernels also apply to the DM kernel. Finally, we provide numerical results that (1) validate the convergence of the heat kernel approximation, (2) demonstrate the greater expressiveness of the DM kernel compared to the Gaussian kernel for supervised learning over a larger class of functions on manifolds with boundary, and (3) demonstrate the advantage of the DM kernel over the Gaussian kernel in learning functions with varying frequencies and co-dimensions.

math.NA

A Geometric Local Parameterization Method for Generalized Hele-Shaw Free Boundary Problems with Source Terms

We develop a meshfree numerical framework for Hele--Shaw free boundary problems with surface tension and source terms based on geometric local parameterization and boundary integral methods. By decomposing the pressure into a particular solution and a harmonic component, the problem is reformulated into a boundary-only system, avoiding volumetric meshing of the evolving domain. For general source terms, we propose an eigenfunction-based approximation on a fixed domain and establish error estimates for both the truncation and coefficient approximation. Numerical experiments verify the accuracy and convergence of the proposed method, and an application to a tumor growth model demonstrates its effectiveness for coupled moving-boundary problems.

math.NA

Learning dynamical systems from noisy data with Weak-form Kernel Ridge Regression

Accurate prediction of complex dynamical systems from noisy measurements remains a significant challenge in scientific computing. Kernel ridge regression learning strategies are often effective when applied to clean data, but have limited success with noisy data. Recent work has observed that a weak formulation can act to filter noisy data, and different learning strategies have achieved increased noise robustness with a weak-form framework. In this manuscript, we give an overview of the filtering mechanism behind the weak formulation and provide a bias-variance error decomposition. Using these insights, we combine a weak formulation with a kernel learning strategy to propose Weak-form Kernel Ridge Regression (WKRR) for learning dynamical systems. The proposed framework is simple to implement, effective for both clean and noisy data, and outperforms several baseline methods. We demonstrate the performance of WKRR on chaotic benchmark systems in up to 64 dimensions, as well as 15,000-dimensional real-world fluid data.

cs.LG

Learning solution operator of dynamical systems with diffusion maps kernel ridge regression

In this work, we propose a simple kernel ridge regression (KRR) framework with a dynamic-aware validation strategy for long-term prediction of complex dynamical systems. By employing a data-driven kernel derived from diffusion maps, the proposed Diffusion Maps Kernel Ridge Regression (DM-KRR) method implicitly adapts to the intrinsic geometry of the system's invariant set, without requiring explicit manifold reconstruction or attractor modeling, procedures that often limit predictive performance. Across a broad range of systems, including smooth manifolds, chaotic attractors, and high-dimensional spatiotemporal flows, DM-KRR consistently outperforms state-of-the-art random feature, neural-network and operator-learning methods in both accuracy and data efficiency. These findings underscore that long-term predictive skill depends not only on model expressiveness, but critically on respecting the geometric constraints encoded in the data through dynamically consistent model selection. Together, simplicity, geometry awareness, and strong empirical performance point to a promising path for reliable and efficient learning of complex dynamical systems.

cs.LG

A model-free method for discovering symmetry in differential equations

Symmetry in differential equations reveals invariances and offers a powerful means to reduce model complexity. Lie group analysis characterizes these symmetries through infinitesimal generators, which provide a local, linear criterion for invariance. However, identifying Lie symmetries directly from scattered data, without explicit knowledge of the governing equations, remains a significant challenge. This work introduces a numerical scheme that approximates infinitesimal generators from data sampled on an unknown smooth manifold, enabling the recovery of continuous symmetries without requiring the analytical form of the differential equations. We employ a manifold learning technique, Generalized Moving Least Squares, to prolongate the data, from which a linear system is constructed whose null space encodes the infinitesimal generators representing the symmetries. Convergence bounds for the proposed approach are derived. Several numerical experiments, including ordinary and partial differential equations, demonstrate the method's accuracy, robustness, and convergence, highlighting its potential for data-driven discovery of symmetries in dynamical systems.

math.NA

A Weak Penalty Neural ODE for Learning Chaotic Dynamics from Noisy Time Series

The accurate forecasting of complex, high-dimensional dynamical systems from observational data is a fundamental task across numerous scientific and engineering disciplines. A significant challenge arises from noisy observations of deterministic dynamics, which severely degrade the performance of data-driven models. In chaotic dynamical systems, where small initial errors amplify exponentially, it is particularly difficult to develop a model from noisy data that achieves short-term accuracy while preserving long-term invariant properties. To overcome this, we consider the weak formulation as a complementary approach to the classical L2-loss function for training models of dynamical systems. We empirically verify that the weak formulation, with a proper choice of test function and integration domain, effectively filters noisy data. This insight explains why a weak form loss function is analogous to fitting a model to filtered data and provides a practical way to parameterize the weak form. Subsequently, we demonstrate how this approach overcomes the instability and inaccuracy of standard Neural ODE (NODE) in modeling chaotic systems. Through numerical examples, we show that our proposed training strategy, the Weak Penalty NODE, is computationally efficient, solver-agnostic, and yields accurate and robust forecasts across benchmark chaotic systems and a real-world climate dataset.

cs.LG

Geometric local parameterization for solving Hele-Shaw problems with surface tension

In this work, we introduce a novel computational framework for solving the two-dimensional Hele-Shaw free boundary problem with surface tension. The moving boundary is represented by point clouds, eliminating the need for a global parameterization. Our approach leverages Generalized Moving Least Squares (GMLS) to construct local geometric charts, enabling high-order approximations of geometric quantities such as curvature directly from the point cloud data. This local parameterization is systematically employed to discretize the governing boundary integral equation, including an analytical formula of the singular integrals. We provide a rigorous convergence analysis for the proposed spatial discretization, establishing consistency and stability under certain conditions. The resulting error bound is derived in terms of the size of the uniformly sampled point cloud data on the moving boundary, the smoothness of the boundary, and the order of the numerical quadrature rule. Numerical experiments confirm the theoretical findings, demonstrating high-order spatial convergence and the expected temporal convergence rates. The method's effectiveness is further illustrated through simulations of complex initial shapes, including interfaces driven by anisotropic surface tension, which correctly evolve towards circular equilibrium states under the influence of surface tension, highlighting the versatility of the method for complex geometry-dependent interface dynamics.

math.NA

Efficient MPC-Based Energy Management System for Secure and Cost-Effective Microgrid Operations

Model predictive control (MPC)-based energy management systems (EMS) are essential for ensuring optimal, secure, and stable operation in microgrids with high penetrations of distributed energy resources. However, due to the high computational cost for the decision-making, the conventional MPC-based EMS typically adopts a simplified integrated-bus power balance model. While this simplification is effective for small networks, large-scale systems require a more detailed branch flow model to account for the increased impact of grid power losses and security constraints. This work proposes an efficient and reliable MPC-based EMS that incorporates power-loss effects and grid-security constraints. %, while adaptively shaping the battery power profile in response to online renewable inputs, achieving reduced operational costs. It enhances system reliability, reduces operational costs, and shows strong potential for online implementation due to its reduced computational effort. Specifically, a second-order cone program (SOCP) branch flow relaxation is integrated into the constraint set, yielding a convex formulation that guarantees globally optimal solutions with high computational efficiency. Owing to the radial topology of the microgrid, this relaxation is practically tight, ensuring equivalence to the original problem. Building on this foundation, an online demand response (DR) module is designed to further reduce the operation cost through peak shaving. To the best of our knowledge, no prior MPC-EMS framework has simultaneously modeled losses and security constraints while coordinating flexible loads within a unified architecture. The developed framework enables secure operation with effective peak shaving and reduced total cost. The effectiveness of the proposed method is validated on 10-bus, 18-bus, and 33-bus systems.

eess.SY

A Higher Order Local Mesh Method for Approximating 1-Laplacians on Unknown Manifolds

We introduce a numerical method for approximating arbitrary differential operators on vector fields in the weak form given point cloud data sampled randomly from a $d$ dimensional manifold embedded in $\mathbb{R}^n$. This method generalizes the local linear mesh method to the local curved mesh method, thus, allowing for the estimation of differential operators with nontrivial Christoffel symbols, such as the Bochner or Hodge Laplacians. In particular, we leverage the potentially small intrinsic dimension of the manifold $(d \ll n)$ to construct local parameterizations that incorporate both local meshes and higher-order curvature information. The former is constructed using low dimensional meshes obtained from local data projected to the tangent spaces, while the latter is obtained by fitting local polynomials with the generalized moving least squares. Theoretically, we prove the spectral convergence for the proposed method for the estimation of the Bochner Laplacian. We provide numerical results supporting the theoretical convergence rates for the Bochner and Hodge Laplacians on simple manifolds.

math.NA

Learning Coarse-Grained Dynamics on Graph

We consider a Graph Neural Network (GNN) non-Markovian modeling framework to identify coarse-grained dynamical systems on graphs. Our main idea is to systematically determine the GNN architecture by inspecting how the leading term of the Mori-Zwanzig memory term depends on the coarse-grained interaction coefficients that encode the graph topology. Based on this analysis, we found that the appropriate GNN architecture that will account for $K$-hop dynamical interactions has to employ a Message Passing (MP) mechanism with at least $2K$ steps. We also deduce that the memory length required for an accurate closure model decreases as a function of the interaction strength under the assumption that the interaction strength exhibits a power law that decays as a function of the hop distance. Supporting numerical demonstrations on two examples, a heterogeneous Kuramoto oscillator model and a power system, suggest that the proposed GNN architecture can predict the coarse-grained dynamics under fixed and time-varying graph topologies.

math.NA

Generalized Finite Difference Method on unknown manifolds

In this paper, we extend the Generalized Finite Difference Method (GFDM) on unknown compact submanifolds of the Euclidean domain, identified by randomly sampled data that (almost surely) lie on the interior of the manifolds. Theoretically, we formalize GFDM by exploiting a representation of smooth functions on the manifolds with Taylor's expansions of polynomials defined on the tangent bundles. We illustrate the approach by approximating the Laplace-Beltrami operator, where a stable approximation is achieved by a combination of Generalized Moving Least-Squares algorithm and novel linear programming that relaxes the diagonal-dominant constraint for the estimator to allow for a feasible solution even when higher-order polynomials are employed. We establish the theoretical convergence of GFDM in solving Poisson PDEs and numerically demonstrate the accuracy on simple smooth manifolds of low and moderate high co-dimensions as well as unknown 2D surfaces. For the Dirichlet Poisson problem where no data points on the boundaries are available, we employ GFDM with the volume-constraint approach that imposes the boundary conditions on data points close to the boundary. When the location of the boundary is unknown, we introduce a novel technique to detect points close to the boundary without needing to estimate the distance of the sampled data points to the boundary. We demonstrate the effectiveness of the volume-constraint employed by imposing the boundary conditions on the data points detected by this new technique compared to imposing the boundary conditions on all points within a certain distance from the boundary, where the latter is sensitive to the choice of truncation distance and require the knowledge of the boundary location.

math.NA

Spectral methods for solving elliptic PDEs on unknown manifolds

In this paper, we propose a mesh-free numerical method for solving elliptic PDEs on unknown manifolds, identified with randomly sampled point cloud data. The PDE solver is formulated as a spectral method where the test function space is the span of the leading eigenfunctions of the Laplacian operator, which are approximated from the point cloud data. While the framework is flexible for any test functional space, we will consider the eigensolutions of a weighted Laplacian obtained from a symmetric Radial Basis Function (RBF) method induced by a weak approximation of a weighted Laplacian on an appropriate Hilbert space. Especially, we consider a test function space that encodes the geometry of the data yet does not require us to identify and use the sampling density of the point cloud. To attain a more accurate approximation of the expansion coefficients, we adopt a second-order tangent space estimation method to improve the RBF interpolation accuracy in estimating the tangential derivatives. This spectral framework allows us to efficiently solve the PDE many times subjected to different parameters, which reduces the computational cost in the related inverse problem applications. In a well-posed elliptic PDE setting with randomly sampled point cloud data, we provide a theoretical analysis to demonstrate the convergent of the proposed solver as the sample size increases. We also report some numerical studies that show the convergence of the spectral solver on simple manifolds and unknown, rough surfaces. Our numerical results suggest that the proposed method is more accurate than a graph Laplacian-based solver on smooth manifolds. On rough manifolds, these two approaches are comparable. Due to the flexibility of the framework, we empirically found improved accuracies in both smoothed and unsmoothed Stanford bunny domains by blending the graph Laplacian eigensolutions and RBF interpolator.

math.NA

A Data-Driven Statistical-Stochastic Surrogate Modeling Strategy for Complex Nonlinear Non-stationary Dynamics

We propose a statistical-stochastic surrogate modeling approach to predict the response of the mean and variance statistics under various initial conditions and external forcing perturbations. The proposed modeling framework extends the purely statistical modeling approach that is practically limited to the homogeneous statistical regime for high-dimensional state variables. The new closure system allows one to overcome several practical issues that emerge in the non-homogeneous statistical regimes. First, the proposed ensemble modeling that couples the mean statistics and stochastic fluctuations naturally produces positive-definite covariance matrix estimation, which is a challenging issue that hampers the purely statistical modeling approaches. Second, the proposed closure model, which embeds a non-Markovian neural-network model for the unresolved fluxes such that the variance of the dynamics is consistent, overcomes the inherent instability of the stochastic fluctuation dynamics. Effectively, the proposed framework extends the classical stochastic parametric modeling paradigm for the unresolved dynamics to a semi-parametric parameterization with a residual Long-Short-Term-Memory neural network architecture. Third, based on empirical information metric, we provide an efficient and effective training procedure by fitting a loss function that measures the differences between response statistics. Supporting numerical examples are provided with the Lorenz-96 model, a system of ODEs that admits the characteristic of chaotic dynamics with both homogeneous and inhomogeneous statistical regimes. In the latter case, we will see the effectiveness of the statistical prediction even though the resolved Fourier modes corresponding to the leading mean energy and variance spectra do not coincide.

physics.data-an

Radial basis approximation of tensor fields on manifolds: From operator estimation to manifold learning

In this paper, we study the Radial Basis Function (RBF) approximation to differential operators on smooth tensor fields defined on closed Riemannian submanifolds of Euclidean space, identified by randomly sampled point cloud data. {The formulation in this paper leverages a fundamental fact that the covariant derivative on a submanifold is the projection of the directional derivative in the ambient Euclidean space onto the tangent space of the submanifold. To differentiate a test function (or vector field) on the submanifold with respect to the Euclidean metric, the RBF interpolation is applied to extend the function (or vector field) in the ambient Euclidean space. When the manifolds are unknown, we develop an improved second-order local SVD technique for estimating local tangent spaces on the manifold. When the classical pointwise non-symmetric RBF formulation is used to solve Laplacian eigenvalue problems, we found that while accurate estimation of the leading spectra can be obtained with large enough data, such an approximation often produces irrelevant complex-valued spectra (or pollution) as the true spectra are real-valued and positive. To avoid such an issue,} we introduce a symmetric RBF discrete approximation of the Laplacians induced by a weak formulation on appropriate Hilbert spaces. Unlike the non-symmetric approximation, this formulation guarantees non-negative real-valued spectra and the orthogonality of the eigenvectors. Theoretically, we establish the convergence of the eigenpairs of both the Laplace-Beltrami operator and Bochner Laplacian {for the symmetric formulation} in the limit of large data with convergence rates. Numerically, we provide supporting examples for approximations of the Laplace-Beltrami operator and various vector Laplacians, including the Bochner, Hodge, and Lichnerowicz Laplacians.

math.NA

Spectral Convergence of Symmetrized Graph Laplacian on manifolds with boundary

We study the spectral convergence of a symmetrized Graph Laplacian matrix induced by a Gaussian kernel evaluated on pairs of embedded data, sampled from a manifold with boundary, a sub-manifold of $\mathbb{R}^m$. Specifically, we deduce the convergence rates for eigenpairs of the discrete Graph-Laplacian matrix to the eigensolutions of the Laplace-Beltrami operator that are well-defined on manifolds with boundary, including the homogeneous Neumann and Dirichlet boundary conditions. For the Dirichlet problem, we deduce the convergence of the \emph{truncated Graph Laplacian}, which is recently numerically observed in applications, and provide a detailed numerical investigation on simple manifolds. Our method of proof relies on the min-max argument over a compact and symmetric integral operator, leveraging the RKHS theory for spectral convergence of integral operator and a recent pointwise asymptotic result of a Gaussian kernel integral operator on manifolds with boundary.

math.NA

Stationary Density Estimation of Itô Diffusions Using Deep Learning

In this paper, we consider the density estimation problem associated with the stationary measure of ergodic Itô diffusions from a discrete-time series that approximate the solutions of the stochastic differential equations. To take an advantage of the characterization of density function through the stationary solution of a parabolic-type Fokker-Planck PDE, we proceed as follows. First, we employ deep neural networks to approximate the drift and diffusion terms of the SDE by solving appropriate supervised learning tasks. Subsequently, we solve a steady-state Fokker-Plank equation associated with the estimated drift and diffusion coefficients with a neural-network-based least-squares method. We establish the convergence of the proposed scheme under appropriate mathematical assumptions, accounting for the generalization errors induced by regressing the drift and diffusion coefficients, and the PDE solvers. This theoretical study relies on a recent perturbation theory of Markov chain result that shows a linear dependence of the density estimation to the error in estimating the drift term, and generalization error results of nonparametric regression and of PDE regression solution obtained with neural-network models. The effectiveness of this method is reflected by numerical simulations of a two-dimensional Student's t distribution and a 20-dimensional Langevin dynamics.

math.NA

Machine Learning-Based Statistical Closure Models for Turbulent Dynamical Systems

We propose a Machine Learning (ML) non-Markovian closure modeling framework for accurate predictions of statistical responses of turbulent dynamical systems subjected to external forcings. One of the difficulties in this statistical closure problem is the lack of training data, which is a configuration that is not desirable in supervised learning with neural network models. In this study with the 40-dimensional Lorenz-96 model, the shortage of data (in temporal) is due to the stationarity of the statistics beyond the decorrelation time, thus, the only informative content in the training data is on short-time transient statistics. We adopted a unified closure framework on various truncation regimes, including and excluding the detailed dynamical equations for the variances. The closure frameworks employ a Long-Short-Term-Memory architecture to represent the higher-order unresolved statistical feedbacks with careful consideration to account for the intrinsic instability yet producing stable long-time predictions. We found that this unified agnostic ML approach performs well under various truncation scenarios. Numerically, the ML closure model can accurately predict the long-time statistical responses subjected to various time-dependent external forces that are not (and maximum forcing amplitudes that are relatively larger than those) in the training dataset.

physics.comp-ph

A Bayesian Machine Learning Algorithm for Predicting ENSO Using Short Observational Time Series

A simple and efficient Bayesian machine learning (BML) training and forecasting algorithm, which exploits only a 20-year short observational time series and an approximate prior model, is developed to predict the Niño 3 sea surface temperature (SST) index. The BML forecast significantly outperforms model-based ensemble predictions and standard machine learning forecasts. Even with a simple feedforward neural network, the BML forecast is skillful for 9.5 months. Remarkably, the BML forecast overcomes the spring predictability barrier to a large extent: the forecast starting from spring remains skillful for nearly 10 months. The BML algorithm can also effectively utilize multiscale features: the BML forecast of SST using SST, thermocline, and wind burst improves on the BML forecast using just SST by at least 2 months. Finally, the BML algorithm also reduces the forecast uncertainty of neural networks and is robust to input perturbations.

physics.ao-ph