SearcharxivSearch

arXiv subjects

Shuhao Cao

Publications and source records attributed to Shuhao Cao.

At least 19 recordsLinked to original sources

Advanced Long-term Earth System Forecasting

Reliable long-term forecasting of Earth system dynamics is fundamentally limited by instabilities in current artificial intelligence (AI) models during extended autoregressive simulations. These failures often originate from inherent spectral bias, leading to inadequate representation of critical high-frequency, small-scale processes and subsequent uncontrolled error amplification. Inspired by the nested grids in numerical models used to resolve small scales, we present TritonCast. At the core of its design is a dedicated latent dynamical core, which ensures the long-term stability of the macro-evolution at a coarse scale. An outer structure then fuses this stable trend with fine-grained local details. This design effectively mitigates the spectral bias caused by cross-scale interactions. In atmospheric science, it achieves state-of-the-art accuracy on the WeatherBench 2 benchmark while demonstrating exceptional long-term stability: executing year-long autoregressive global forecasts and completing multi-year climate simulations that span the entire available $2500$-day test period without drift. In oceanography, it extends skillful eddy forecast to $120$ days and exhibits unprecedented zero-shot cross-resolution generalization. Ablation studies reveal that this performance stems from the synergistic interplay of the architecture's core components. TritonCast thus offers a promising pathway towards a new generation of trustworthy, AI-driven simulations. This significant advance has the potential to accelerate discovery in climate and Earth system science, enabling more reliable long-term forecasting and deeper insights into complex geophysical dynamics.

cs.LG

Solver-in-the-Loop joint operator learning: fractional Laplace-Beltrami features for interface reconstruction

In this work, we propose a joint operator learning method for reconstructing images of conductivity coefficients from boundary data. Inspired by the idea of employing partial differential equation (PDE) solvers as preconditioners for this inverse problem, we investigate a ``solver-in-the-loop'' training mechanism. It allows the interaction of learnable parameters integrated in a PDE solver module and those in neural networks for reconstructing images. Specifically, we employ a fractional Laplace-Beltrami operator with a learnable fractional order, which transforms boundary data into high-dimensional features. These features then serve as input to a neural network, significantly improving reconstruction accuracy. For this purpose, a Learning-Automated FEM (LA-FEM) package, facilitating this ``solver-in-the-loop'' property, is developed with PyTorch as a backend. The new LA-FEM module conveniently allows the auto-differentiation regarding an objective function to freely propagate through the PDE solver from the forward problem and the coupled neural networks for the inverse problem.

math.NA

Mitigating spectral bias for the multiscale operator learning

Neural operators have emerged as a powerful tool for learning the mapping between infinite-dimensional parameter and solution spaces of partial differential equations (PDEs). In this work, we focus on multiscale PDEs that have important applications such as reservoir modeling and turbulence prediction. We demonstrate that for such PDEs, the spectral bias towards low-frequency components presents a significant challenge for existing neural operators. To address this challenge, we propose a hierarchical attention neural operator (HANO) inspired by the hierarchical matrix approach. HANO features a scale-adaptive interaction range and self-attentions over a hierarchy of levels, enabling nested feature computation with controllable linear cost and encoding/decoding of multiscale solution space. We also incorporate an empirical $H^1$ loss function to enhance the learning of high-frequency components. Our numerical experiments demonstrate that HANO outperforms state-of-the-art (SOTA) methods for representative multiscale problems.

cs.LG

Spectral-Refiner: Accurate Fine-Tuning of Spatiotemporal Fourier Neural Operator for Turbulent Flows

Recent advancements in operator-type neural networks have shown promising results in approximating the solutions of spatiotemporal Partial Differential Equations (PDEs). However, these neural networks often entail considerable training expenses, and may not always achieve the desired accuracy required in many scientific and engineering disciplines. In this paper, we propose a new learning framework to address these issues. A new spatiotemporal adaptation is proposed to generalize any Fourier Neural Operator (FNO) variant to learn maps between Bochner spaces, which can perform an arbitrary-length temporal super-resolution for the first time. To better exploit this capacity, a new paradigm is proposed to refine the commonly adopted end-to-end neural operator training and evaluations with the help from the wisdom from traditional numerical PDE theory and techniques. Specifically, in the learning problems for the turbulent flow modeled by the Navier-Stokes Equations (NSE), the proposed paradigm trains an FNO only for a few epochs. Then, only the newly proposed spatiotemporal spectral convolution layer is fine-tuned without the frequency truncation. The spectral fine-tuning loss function uses a negative Sobolev norm for the first time in operator learning, defined through a reliable functional-type a posteriori error estimator whose evaluation is exact thanks to the Parseval identity. Moreover, unlike the difficult nonconvex optimization problems in the end-to-end training, this fine-tuning loss is convex. Numerical experiments on commonly used NSE benchmarks demonstrate significant improvements in both computational efficiency and accuracy, compared to end-to-end evaluation and traditional numerical PDE solvers under certain conditions. The source code is publicly available at https://github.com/scaomath/torch-cfd.

cs.LG

Edge-averaged virtual element methods for convection-diffusion and convection-dominated problems

This manuscript develops edge-averaged virtual element (EAVE) methodologies to address convection-diffusion problems effectively in the convection-dominated regime. It introduces a variant of EAVE that ensures monotonicity (producing an $M$-matrix) on Voronoi polygonal meshes, provided their duals are Delaunay triangulations with acute angles. Furthermore, the study outlines a comprehensive framework for EAVE methodologies, introducing another variant that integrates with the stiffness matrix derived from the lowest-order virtual element method for the Poisson equation. Numerical experiments confirm the theoretical advantages of the monotonicity property and demonstrate an optimal convergence rate across various mesh configurations.

math.NA

A numerical domain decomposition method for solving elliptic equations on manifolds

A new numerical domain decomposition method is proposed for solving elliptic equations on compact Riemannian manifolds. The advantage of this method is to avoid global triangulations or grids on manifolds. Our method is numerically tested on some $4$-dimensional manifolds such as the unit sphere $S^{4}$, the complex projective space $\mathbb{CP}^{2}$ and the product manifold $S^{2} \times S^{2}$.

math.NA

Convergence and optimality of an adaptive modified weak Galerkin finite element method

An adaptive modified weak Galerkin method (AmWG) for an elliptic problem is studied in this paper, in addition to its convergence and optimality. The modified weak Galerkin bilinear form is simplified without the need of the skeletal variable, and the approximation space is chosen as the discontinuous polynomial space as in the discontinuous Galerkin method. Upon a reliable residual-based a posteriori error estimator, an adaptive algorithm is proposed together with its convergence and quasi-optimality proved for the lowest order case. The primary tool is to bridge the connection between the modified weak Galerkin method and the Crouzeix-Raviart nonconforming finite element. Unlike the traditional convergence analysis for methods with a discontinuous polynomial approximation space, the convergence of AmWG is penalty parameter free. Numerical results are presented to support the theoretical results.

math.NA

Immersed Virtual Element Methods for Electromagnetic Interface Problems in Three Dimensions

Finite element methods for electromagnetic problems modeled by Maxwell-type equations are highly sensitive to the conformity of approximation spaces, and non-conforming methods may cause loss of convergence. This fact leads to an essential obstacle for almost all the interface-unfitted mesh methods in the literature regarding the application to electromagnetic interface problems, as they are based on non-conforming spaces. In this work, a novel immersed virtual element method for solving a 3D $\mathbf{H}(\mathrm{curl})$ interface problem is developed, and the motivation is to combine the conformity of virtual element spaces and robust approximation capabilities of immersed finite element spaces. The proposed method is able to achieve optimal convergence. To develop a systematic framework, the $H^1$, $\mathbf{H}(\mathrm{curl})$ and $\mathbf{H}(\mathrm{div})$ interface problems and their corresponding problem-orientated immersed virtual element spaces are considered all together. In addition, the de Rham complex will be established based on which the Hiptmair-Xu (HX) preconditioner can be used to develop a fast solver for the $\mathbf{H}(\mathrm{curl})$ interface problem.

math.NA

Transformer Meets Boundary Value Inverse Problems

A Transformer-based deep direct sampling method is proposed for electrical impedance tomography, a well-known severely ill-posed nonlinear boundary value inverse problem. A real-time reconstruction is achieved by evaluating the learned inverse operator between carefully designed data and the reconstructed images. An effort is made to give a specific example to a fundamental question: whether and how one can benefit from the theoretical structure of a mathematical problem to develop task-oriented and structure-conforming deep neural networks? Specifically, inspired by direct sampling methods for inverse problems, the 1D boundary data in different frequencies are preprocessed by a partial differential equation-based feature map to yield 2D harmonic extensions as different input channels. Then, by introducing learnable non-local kernels, the direct sampling is recast to a modified attention mechanism. The new method achieves superior accuracy over its predecessors and contemporary operator learners and shows robustness to noises in benchmarks. This research shall strengthen the insights that, despite being invented for natural language processing tasks, the attention mechanism offers great flexibility to be modified in conformity with the a priori mathematical knowledge, which ultimately leads to the design of more physics-compatible neural architectures.

cs.LG

A Bootstrap Multigrid Eigensolver

This paper introduces bootstrap multigrid methods for solving eigenvalue problems arising from the discretization of partial differential equations. Inspired by the full bootstrap algebraic multigrid (BAMG) setup algorithm that includes an AMG eigensolver, it is illustrated how the algorithm can be simplified for the case of a discretized partial differential equation (PDE), thereby developing a bootstrap geometric multigrid (BMG) approach. We illustrate numerically the efficacy of the BMG method for: (1) recovering eigenvalues having large multiplicity, (2) computing interior eigenvalues, and (3) approximating shifted indefinite eigenvalue problems. Numerical experiments are presented to illustrate the basic components and ideas behind the success of the overall bootstrap multigrid approach. For completeness, we present a simplified error analysis of a two-grid bootstrap algorithm for the Laplace-Beltrami eigenvalue problem.

math.NA

Immersed Virtual Element Methods for Elliptic Interface Problems in Two Dimensions

This article presents an immersed virtual element method for solving a class of interface problems that combines the advantages of both body-fitted mesh methods and unfitted mesh methods. A background body-fitted mesh is generated initially. On those interface elements, virtual element spaces are constructed as solution spaces to local interface problems, and exact sequences can be established for these new spaces involving discontinuous coefficients. The discontinuous coefficients of interface problems are recast as Hodge star operators that are the key to project immersed virtual functions to classic immersed finite element (IFE) functions for computing numerical solutions. An a priori convergence analysis is established robust with respect to the interface location. The proposed method is capable of handling more complicated interface element configuration and provides better performance than the conventional penalty-type IFE method for the H(curl)-interface problem arising from Maxwell equations. It also brings a connection between various methods such as body-fitted methods, IFE methods, virtual element methods, etc.

math.NA

A nonconforming primal hybrid finite element method for the two-dimensional vector Laplacian

We introduce a nonconforming hybrid finite element method for the two-dimensional vector Laplacian, based on a primal variational principle for which conforming methods are known to be inconsistent. Consistency is ensured using penalty terms similar to those used to stabilize hybridizable discontinuous Galerkin (HDG) methods, with a carefully chosen penalty parameter due to Brenner, Li, and Sung [Math. Comp., 76 (2007), pp. 573-595]. Our method accommodates elements of arbitrarily high order and, like HDG methods, it may be implemented efficiently using static condensation. The lowest-order case recovers the $P_1$-nonconforming method of Brenner, Cui, Li, and Sung [Numer. Math., 109 (2008), pp. 509-533], and we show that higher-order convergence is achieved under appropriate regularity assumptions. The analysis makes novel use of a family of weighted Sobolev spaces, due to Kondrat'ev, for domains admitting corner singularities.

math.NA

A Virtual Finite Element Method for Two Dimensional Maxwell Interface Problems with a Background Unfitted Mesh

A virtual element method (VEM) with the first order optimal convergence order is developed for solving two-dimensional Maxwell interface problems on a special class of polygonal meshes that are cut by the interface from a background unfitted mesh. A novel virtual space is introduced on a virtual triangulation of the polygonal mesh satisfying a maximum angle condition, which shares exactly the same degrees of freedom as the usual H(curl)-conforming virtual space. This new virtual space serves as the key to prove that the optimal error bounds of the VEM are independent of high aspect ratio of the possible anisotropic polygonal mesh near the interface.

math.NA

How to Understand Masked Autoencoders

"Masked Autoencoders (MAE) Are Scalable Vision Learners" revolutionizes the self-supervised learning method in that it not only achieves the state-of-the-art for image pre-training, but is also a milestone that bridges the gap between visual and linguistic masked autoencoding (BERT-style) pre-trainings. However, to our knowledge, to date there are no theoretical perspectives to explain the powerful expressivity of MAE. In this paper, we, for the first time, propose a unified theoretical framework that provides a mathematical understanding for MAE. Specifically, we explain the patch-based attention approaches of MAE using an integral kernel under a non-overlapping domain decomposition setting. To help the research community to further comprehend the main reasons of the great success of MAE, based on our framework, we pose five questions and answer them with mathematical rigor using insights from operator theory.

cs.CV

Error analysis of a decoupled finite element method for quad-curl problems

Finite element approximation to a decoupled formulation for the quad--curl problem is studied in this paper. The difficulty of constructing elements with certain conformity to the quad--curl problems has been greatly reduced. For convex domains, where the regularity assumption holds for Stokes equation, the approximation to the curl of the true solution has quadratic order of convergence and first order for the energy norm. If the solution shows singularity, an a posterior error estimator is developed and a separate marking adaptive finite element procedure is proposed, together with its convergence proved. Both the a priori and a posteriori error analysis are supported by the numerical examples.

math.NA

Choose a Transformer: Fourier or Galerkin

In this paper, we apply the self-attention from the state-of-the-art Transformer in Attention Is All You Need for the first time to a data-driven operator learning problem related to partial differential equations. An effort is put together to explain the heuristics of, and to improve the efficacy of the attention mechanism. By employing the operator approximation theory in Hilbert spaces, it is demonstrated for the first time that the softmax normalization in the scaled dot-product attention is sufficient but not necessary. Without softmax, the approximation capacity of a linearized Transformer variant can be proved to be comparable to a Petrov-Galerkin projection layer-wise, and the estimate is independent with respect to the sequence length. A new layer normalization scheme mimicking the Petrov-Galerkin projection is proposed to allow a scaling to propagate through attention layers, which helps the model achieve remarkable accuracy in operator learning tasks with unnormalized data. Finally, we present three operator learning experiments, including the viscid Burgers' equation, an interface Darcy flow, and an inverse interface coefficient identification problem. The newly proposed simple attention-based operator learner, Galerkin Transformer, shows significant improvements in both training cost and evaluation accuracy over its softmax-normalized counterparts.

cs.LG

Neural networks with trainable matrix activation functions

The training process of neural networks usually optimize weights and bias parameters of linear transformations, while nonlinear activation functions are pre-specified and fixed. This work develops a systematic approach to constructing matrix-valued activation functions whose entries are generalized from ReLU. The activation is based on matrix-vector multiplications using only scalar multiplications and comparisons. The proposed activation functions depend on parameters that are trained along with the weights and bias vectors. Neural networks based on this approach are simple and efficient and are shown to be robust in numerical experiments.

cs.LG

A simple virtual element-based flux recovery on quadtree

In this paper, we introduce a simple local flux recovery for $\mathcal{Q}_k$ finite element of a scalar coefficient diffusion equation on quadtree meshes, with no restriction on the irregularities of hanging nodes. The construction requires no specific ad hoc tweaking for hanging nodes on $l$-irregular ($l\geq 2$) meshes thanks to the adoption of virtual element families. The rectangular elements with hanging nodes are treated as polygons as in the flux recovery context. An efficient a posteriori error estimator is then constructed based on the recovered flux, and its reliability is proved under common assumptions, both of which are further verified in numerics.

math.NA