Searcharxiv⌕ Search

arXiv subjects

Stéphane Gaudreault

Publications and source records attributed to Stéphane Gaudreault.

11 recordsLinked to original sources

Learning to Advect: A Neural Semi-Lagrangian Architecture for Weather Forecasting

Machine-learning approaches to weather forecasting often employ a monolithic architecture in which distinct physical mechanisms, such as advection, diffusive mixing, thermodynamic processes, and forcing, are represented implicitly within a single large neural network. This is particularly problematic for advection, where long-range transport typically requires expensive global interaction mechanisms or deep stacks of local convolutional layers. To address this limitation, we introduce a physics-inspired neural architecture that decomposes latent-state evolution into dedicated advection, diffusion, and reaction operators. Its central component is a Neural Semi-Lagrangian operator that performs trajectory-based transport via differentiable interpolation on the sphere, allowing the network to learn both a compressed set of latent modes to be transported and their characteristic trajectories. The atmospheric state is projected into latent space and spatially coarsened to a processor grid, where advection, diffusion, and reaction operators jointly evolve the representation. Diffusive mixing and unresolved dissipation are represented by depthwise-separable spatial mixing, while local source terms and vertical interactions are handled through pointwise channel interactions. We evaluate a reference implementation of the proposed architecture on global weather forecasting. Evaluated on ERA5 benchmarks, the reference model achieves competitive deterministic forecast skill, with particularly strong performance at short to medium lead times, while preserving improved spectral fidelity and forecast activity relative to several leading data-driven baselines.

cs.LG↗

Thermodynamically Constrained Information Geometric Regularization for Compressible Flows

We construct and analyze a thermodynamic extension of the recently proposed information geometric regularization of Cao and Schäfer. The construction extends their shock-mitigating Hessian metric geometry using the Shannon entropy to constrain the regularized motion based on a thermodynamic length. Reformulating the equations in terms of mass and specific entropy explicitly connects the thermodynamic state to a position in the diffeomorphism group, allowing for a derivation of the regularized equations using an information geometric mechanics formalism based on geodesics on a Hessian manifold with a dual affine connection. The dynamics are defined using a pullback geometry for the Levi--Civita connection, describing constrained geodesic motion, and the cubic Amari--Chentsov tensor describing the information geometric correction. This new compressible fluid model introduces an anisotropic stress tensor to the momentum equation that vanishes along isentropic directions and an additional elliptic equation coupled to the barotropic regularization. Numerical simulations in one and two spatial dimensions demonstrate that the geometrically consistent incorporation of a thermodynamic constraint mitigates cusp singularities previously observed in other approaches while still maintaining the benefits of an inviscid regularization.

math.NA↗

Characteristic Bending in Incompressible Flows

We present the Characteristic Bending (CB) method, a general framework for advecting quantities under incompressible velocity fields. The method builds on standard semi-Lagrangian advection by interpreting the backward-in-time characteristic reconstruction as the construction of a reference map, a diffeomorphism between the current and initial geometries of the advected space. From this viewpoint, the CB method applies a volume-preserving projection to the map, systematically removing spurious compressible errors arising from time integration, interpolation, or from velocity fields that are only approximately divergence-free. This projection bends the characteristics toward the divergence-free space, preserving mass and geometric features of the advected fields, even in the presence of significant error. We demonstrate the method in both two and three dimensions using benchmark problems and for multiphase flows governed by the incompressible Navier-Stokes equations. The results show that the CB method serves as a drop-in replacement for traditional semi-Lagrangian schemes and as an augmentation of reference map formulations, offering improved robustness and accuracy in incompressible flow simulations.

physics.flu-dyn↗

DEF: Diffusion-augmented Ensemble Forecasting

We present DEF (\textbf{\ul{D}}iffusion-augmented \textbf{\ul{E}}nsemble \textbf{\ul{F}}orecasting), a novel approach for generating initial condition perturbations. Modern approaches to initial condition perturbations are primarily designed for numerical weather prediction (NWP) solvers, limiting their applicability in the rapidly growing field of machine learning for weather prediction. Consequently, stochastic models in this domain are often developed on a case-by-case basis. We demonstrate that a simple conditional diffusion model can (1) generate meaningful structured perturbations, (2) be applied iteratively, and (3) utilize a guidance term to intuitivey control the level of perturbation. This method enables the transformation of any deterministic neural forecasting system into a stochastic one. With our stochastic extended systems, we show that the model accumulates less error over long-term forecasts while producing meaningful forecast distributions. We validate our approach on the 5.625$^\circ$ ERA5 reanalysis dataset, which comprises atmospheric and surface variables over a discretized global grid, spanning from the 1960s to the present. On this dataset, our method demonstrates improved predictive performance along with reasonable spread estimates.

cs.LG↗

Probabilistic Forecasting for Dynamical Systems with Missing or Imperfect Data

The modeling of dynamical systems is essential in many fields, but applying machine learning techniques is often challenging due to incomplete or noisy data. This study introduces a variant of stochastic interpolation (SI) for probabilistic forecasting, estimating future states as distributions rather than single-point predictions. We explore its mathematical foundations and demonstrate its effectiveness on various dynamical systems, including the challenging WeatherBench dataset.

physics.comp-ph↗

PEARL: Preconditioner Enhancement through Actor-critic Reinforcement Learning

We present PEARL (Preconditioner Enhancement through Actor-critic Reinforcement Learning), a novel approach to learning matrix preconditioners. Existing preconditioners such as Jacobi, Incomplete LU, and Algebraic Multigrid methods offer problem-specific advantages but rely heavily on hyperparameter tuning. Recent advances have explored using deep neural networks to learn preconditioners, though challenges such as misbehaved objective functions and costly training procedures remain. PEARL introduces a reinforcement learning approach for learning preconditioners, specifically, a contextual bandit formulation. The framework utilizes an actor-critic model, where the actor generates the incomplete Cholesky decomposition of preconditioners, and the critic evaluates them based on reward-specific feedback. To further guide the training, we design a dual-objective function, combining updates from the critic and condition number. PEARL contributes a generalizable preconditioner learning method, dynamic sparsity exploration, and cosine schedulers for improved stability and exploratory power. We compare our approach to traditional and neural preconditioners, demonstrating improved flexibility and iterative solving speed.

cs.LG↗

Low-synchronization Arnoldi Methods for the Matrix Exponential with Application to Exponential Integrators

High order exponential integrators require computing linear combination of exponential like $φ$-functions of large matrices $A$ times a vector $v$. Krylov projection methods are the most general and remain an efficient choice for computing the matrix-function-vector-product evaluation when the matrix is $A$ is large and unable to be explicitly stored, or when obtaining information about the spectrum is expensive. The Krylov approximation relies on the Gram-Schmidt (GS) orthogonalization procedure to produce the orthonormal basis $V_m$. In parallel, GS orthogonalization requires \textit{global synchronizations} for inner products and vector normalization in the orthogonalization process. Reducing the amount of global synchronizations is of paramount importance for the efficiency of a numerical algorithm in a massively parallel setting. We improve the parallel strong scaling properties of exponential integrators by addressing the underlying bottleneck in the linear algebra using low-synchronization GS methods. The resulting orthogonalization algorithms have an accuracy comparable to modified Gram-Schmidt yet are better suited for distributed architecture, as only one global communication is required per orthogonalization-step. We present geophysics-based numerical experiments and standard examples routinely used to test stiff time integrators, which validate that reducing global communication leads to better parallel scalability and reduced time-to-solution for exponential integrators.

math.NA↗

Deep Learning for Koopman Operator Estimation in Idealized Atmospheric Dynamics

Deep learning is revolutionizing weather forecasting, with new data-driven models achieving accuracy on par with operational physical models for medium-term predictions. However, these models often lack interpretability, making their underlying dynamics difficult to understand and explain. This paper proposes methodologies to estimate the Koopman operator, providing a linear representation of complex nonlinear dynamics to enhance the transparency of data-driven models. Despite its potential, applying the Koopman operator to large-scale problems, such as atmospheric modeling, remains challenging. This study aims to identify the limitations of existing methods, refine these models to overcome various bottlenecks, and introduce novel convolutional neural network architectures that capture simplified dynamics.

cs.LG↗

Challenges in computing matrix functions

This manuscript summarizes the outcome of the focus groups at "The f(A)bulous workshop on matrix functions and exponential integrators", held at the Max Planck Institute for Dynamics of Complex Technical Systems in Magdeburg, Germany, on 25-27 September 2023. There were three focus groups in total, each with a different theme: knowledge transfer, high-performance and energy-aware computing, and benchmarking. We collect insights, open issues, and perspectives from each focus group, as well as from general discussions throughout the workshop. Our primary aim is to highlight ripe research directions and continue to build on the momentum from a lively meeting.

math.NA↗

KIOPS: A fast adaptive Krylov subspace solver for exponential integrators

This paper presents a new algorithm KIOPS for computing linear combinations of $φ$-functions that appear in exponential integrators. This algorithm is suitable for large-scale problems in computational physics where little or no information about the spectrum or norm of the Jacobian matrix is known a priori. We first show that such problems can be solved efficiently by computing a single exponential of a modified matrix. Then our approach is to compute an appropriate basis for the Krylov subspace using the incomplete orthogonalization procedure and project the matrix exponential on this subspace. We also present a novel adaptive procedure that significantly reduces the computational complexity of exponential integrators. Our numerical experiments demonstrate that KIOPS outperforms the current state-of-the-art adaptive Krylov algorithm phipm.

math.NA↗

High-order numerical solutions to the shallow-water equations on the rotated cubed-sphere grid

A novel numerical approach to solving the shallow-water equations on the sphere using high-order numerical discretizations in both space and time is proposed. A space-time tensor formalism is used to express the equations of motion covariantly and to describe the geometry of the rotated cubed-sphere grid. The spatial discretization is done with the direct flux reconstruction method, which is an alternative formulation to the discontinuous Galerkin approach. The equations of motion are solved in differential form and the resulting discretization is free from quadrature rules. It is well known that the time step of traditional explicit methods is limited by the phase velocity of the fastest waves. Exponential integration is employed to enable integrations with significantly larger time step sizes and improve the efficiency of the overall time integration. New multistep-type exponential propagation iterative methods of orders 4, 5 and 6 are constructed and applied to integrate the shallow-water equations in time. These new schemes enable time integration with high-order accuracy but without significant increases in computational time compared to low-order methods. The exponential matrix functions-vector products used in the exponential schemes are approximated using the complex-step approximation of the Jacobian in the Krylov-based KIOPS (Krylov with incomplete orthogonalization procedure solver) algorithm. Performance of the new numerical methods is evaluated using a set of standard benchmark tests.

math.NA↗