SearcharxivSearch

arXiv subjects

Ben Southworth

Publications and source records attributed to Ben Southworth.

5 recordsLinked to original sources

On the Convergence Behavior of Preconditioned Gradient Descent Toward the Rich Learning Regime

Spectral bias, the tendency of neural networks to learn low frequencies first, can be both a blessing and a curse. While it enhances the generalization capabilities by suppressing high-frequency noise, it can be a limitation in scientific tasks that require capturing fine-scale structures. The delayed generalization phenomenon known as grokking is another barrier to rapid training of neural networks. Grokking has been hypothesized to arise as learning transitions from the NTK to the feature-rich regime. This paper explores the impact of preconditioned gradient descent (PGD), such as Gauss-Newton, on spectral bias and grokking phenomena. We demonstrate through theoretical and empirical results how PGD can mitigate issues associated with spectral bias. Additionally, building on the rich learning regime grokking hypothesis, we study how PGD can be used to reduce delays associated with grokking. Our conjecture is that PGD, without the impediment of spectral bias, enables uniform exploration of the parameter space in the NTK regime. Our experimental results confirm this prediction, providing strong evidence that grokking represents a transitional behavior between the lazy regime characterized by the NTK and the rich regime. These findings deepen our understanding of the interplay between optimization dynamics, spectral bias, and the phases of neural network learning.

cs.LG

Learning interpretable closures for thermal radiation transport in optically-thin media using WSINDy

We introduce an equation learning framework to identify a closed set of equations for moment quantities in 1D thermal radiation transport (TRT) in optically thin media. While optically thick media admits a well-known diffusive closure, the utility of moment closures in providing accurate low-dimensional surrogates for TRT in optically thin media is unclear, as the mean-free path of photons is large and the radiation flux is far from its Fickean limit. Here, we demonstrate the viability of using weak-form equation learning to close the system of equations for the energy density, radiation flux, and temperature in optically thin TRT. We show that the WSINDy algorithm (Weak-form Sparse Identification of Nonlinear Dynamics), together with an advantageous change of variables and an auxiliary equation for the radiation-energy-weighted opacity, enables robust and efficient identification of closures that preserve many desired physical properties from the high fidelity system, including hyperbolicity, rotational symmetry, black-body equilibria, and linear stability of black-body equilibria, all of which manifest as library constraints or convex constraints on the closure coefficients. Crucially, the weak form enables closures to be learned from simulation data with ray effects and particle noise, which then do not appear in simulations of the resulting closed moment system. Finally, we demonstrate that our closure models can be extrapolated in the key system parameters of drive temperature $T_{in}$ and scalar opacity $γ$, and this extrapolation is to an extent quantifiable by a Knudsen-like dimensionless parameter.

math.DS

Leveraging KANs for Expedient Training of Multichannel MLPs via Preconditioning and Geometric Refinement

Multilayer perceptrons (MLPs) are a workhorse machine learning architecture, used in a variety of modern deep learning frameworks. However, recently Kolmogorov-Arnold Networks (KANs) have become increasingly popular due to their success on a range of problems, particularly for scientific machine learning tasks. In this paper, we exploit the relationship between KANs and multichannel MLPs to gain structural insight into how to train MLPs faster. We demonstrate the KAN basis (1) provides geometric localized support, and (2) acts as a preconditioned descent in the ReLU basis, overall resulting in expedited training and improved accuracy. Our results show the equivalence between free-knot spline KAN architectures, and a class of MLPs that are refined geometrically along the channel dimension of each weight tensor. We exploit this structural equivalence to define a hierarchical refinement scheme that dramatically accelerates training of the multi-channel MLP architecture. We show further accuracy improvements can be had by allowing the $1$D locations of the spline knots to be trained simultaneously with the weights. These advances are demonstrated on a range of benchmark examples for regression and scientific machine learning.

cs.LG

Asynchronous Truncated Multigrid-reduction-in-time (AT-MGRIT)

In this paper, we present the new "asynchronous truncated multigrid-reduction-in-time" (AT-MGRIT) algorithm for introducing time parallelism to the solution of discretized time-dependent problems. The new algorithm is based on the multigrid-reduction-in-time (MGRIT) approach, which, in certain settings, is equivalent to another common multilevel parallel-in-time method, Parareal. In contrast to Parareal and MGRIT that both consider a global temporal grid over the entire time interval on the coarsest level, the AT-MGRIT algorithm uses truncated local time grids on the coarsest level, each grid covering certain temporal subintervals. These local grids can be solved completely independent from each other, which reduces the sequential part of the algorithm and, thus, increases parallelism in the method. Here, we study the effect of using truncated local coarse grids on the convergence of the algorithm, both theoretically and numerically, and show, using challenging nonlinear problems, that the new algorithm consistently outperforms classical Parareal/MGRIT in terms of time to solution.

math.NA

Surface Deposition of the Enceladus Plume and the Zenith Angle of Emissions

Since the discovery of an ice particle plume erupting from the south polar terrain on Saturn's moon Enceladus, the geophysical mechanisms driving its activity have been the focus of substantial scientific research. The pattern and deposition rate of plume material on Enceladus' surface is of interest because it provides valuable information about the dynamics of the ice particle ejection as well as the surface erosion. Surface deposition maps derived from numerical plume simulations by Kempf et al. (2010) have been used by various researchers to interpret data obtained by various Cassini instruments. Here, an updated and detailed set of deposition maps is provided based on a deep-source plume model (Schmidt et al., 2008), for the eight ice-particle jets identified in Spitale and Porco (2007), the updated set of jets proposed in Porco et al. (2014), and a contrasting curtain-style plume proposed in Spitale et al. (2015). Methods for computing the surface deposition are detailed, and the structure of surface deposition patterns is shown to be consistent across changes in the production rate and size distribution of the plume. Maps are also provided of the surface deposition structure originating in each of the four Tiger Stripes. Finally, the differing approaches used in Porco et al. (2014) and Spitale et al. (2015) have given rise to a jets vs. curtains controversy regarding the emission structure of the Enceladus plume. Here we simulate each, leading to new insight that, over time, most emissions must be directed relatively orthogonal to the surface because jets "tilted" significantly away from orthogonal lead to surface deposition patterns inconsistent with surface images. Data for maps are available in HDF5 format for a variety of particle sizes at http://impact.colorado.edu/southworth_data.

astro-ph.EP