SearcharxivSearch

arXiv subjects

Lachlan MacDonald

Publications and source records attributed to Lachlan MacDonald.

8 recordsLinked to original sources

SGD at the Edge of Stability: Stochastic Stabilization with Large Learning Rates

Modern deep learning has been shown to operate at the edge of stability, routinely using learning rates far larger than those justified by classical optimization theory. Most prior analyses of the edge of stability phenomenon focus on deterministic gradient descent, leaving the stochastic setting largely unexplored. In this work, we provide sharp convergence guarantees for Stochastic Gradient Descent (SGD) applied to the multiclass cross-entropy loss, for both linear classifiers and two-layer neural networks. We show that the stochasticity of SGD may cause the dynamics to alternate between an edge-of-stability regime that is dominated by curvature-driven oscillations, and a stable regime in which the expected loss decreases at a controlled rate. Despite that, we prove that SGD self-stabilizes the dynamics, ensuring that the iterates return to stability in a fixed number of iterations and allowing convergence in the best-iterate sense even with large learning rates. Experiments validate our theoretical findings and illustrate the benefits of SGD in the large-stepsize regime.

stat.ML

Curvature-Aware Training for Coordinate Networks

Coordinate networks are widely used in computer vision due to their ability to represent signals as compressed, continuous entities. However, training these networks with first-order optimizers can be slow, hindering their use in real-time applications. Recent works have opted for shallow voxel-based representations to achieve faster training, but this sacrifices memory efficiency. This work proposes a solution that leverages second-order optimization methods to significantly reduce training times for coordinate networks while maintaining their compressibility. Experiments demonstrate the effectiveness of this approach on various signal modalities, such as audio, images, videos, shape reconstruction, and neural radiance fields.

cs.CV

On Quantizing Implicit Neural Representations

The role of quantization within implicit/coordinate neural networks is still not fully understood. We note that using a canonical fixed quantization scheme during training produces poor performance at low-rates due to the network weight distributions changing over the course of training. In this work, we show that a non-uniform quantization of neural weights can lead to significant improvements. Specifically, we demonstrate that a clustered quantization enables improved reconstruction. Finally, by characterising a trade-off between quantization and network capacity, we demonstrate that it is possible (while memory inefficient) to reconstruct signals using binary neural networks. We demonstrate our findings experimentally on 2D image reconstruction and 3D radiance fields; and show that simple quantization methods and architecture search can achieve compression of NeRF to less than 16kb with minimal loss in performance (323x smaller than the original NeRF).

cs.CV

How You Start Matters for Generalization

Characterizing the remarkable generalization properties of over-parameterized neural networks remains an open problem. In this paper, we promote a shift of focus towards initialization rather than neural architecture or (stochastic) gradient descent to explain this implicit regularization. Through a Fourier lens, we derive a general result for the spectral bias of neural networks and show that the generalization of neural networks is heavily tied to their initialization. Further, we empirically solidify the developed theoretical insights using practical, deep networks. Finally, we make a case against the controversial flat-minima conjecture and show that Fourier analysis grants a more reliable framework for understanding the generalization of neural networks.

cs.LG

On Regularizing Coordinate-MLPs

We show that typical implicit regularization assumptions for deep neural networks (for regression) do not hold for coordinate-MLPs, a family of MLPs that are now ubiquitous in computer vision for representing high-frequency signals. Lack of such implicit bias disrupts smooth interpolations between training samples, and hampers generalizing across signal regions with different spectra. We investigate this behavior through a Fourier lens and uncover that as the bandwidth of a coordinate-MLP is enhanced, lower frequencies tend to get suppressed unless a suitable prior is provided explicitly. Based on these insights, we propose a simple regularization technique that can mitigate the above problem, which can be incorporated into existing networks without any architectural modifications.

cs.LG

A characteristic map for the holonomy groupoid of a foliation

We prove a generalisation of Bott's vanishing theorem for the full transverse frame holonomy groupoid of any transversely orientable foliated manifold. As a consequence we obtain a characteristic map encoding both primary and secondary characteristic classes. Previous descriptions of this characteristic map are formulated for the Morita equivalent étale groupoid obtained via a choice of complete transversal. By working with the full holonomy groupoid we obtain novel geometric representatives of characteristic classes. In particular we give a geometric, non-étale analogue of the codimension 1 Godbillon-Vey cyclic cocycle of Connes and Moscovici in terms of path integrals of the curvature form of a Bott connection.

math.DG

Equivariant $KK$-theory for non-Hausdorff groupoids

We give a detailed and unified survey of equivariant $KK$-theory over locally compact, second countable, locally Hausdorff groupoids. We indicate precisely how the "classical" proofs relating to the Kasparov product can be used almost word-for-word in this setting, and give proofs for several results which do not currently appear in the literature.

math.KT

The Godbillon-Vey invariant in equivariant $KK$-theory

We construct a groupoid equivariant Kasparov class for transversely oriented foliations in all codimensions. In codimension 1 we show that the Chern character of an associated semifinite spectral triple recovers the Connes-Moscovici cyclic cocycle for the Godbillon-Vey secondary characteristic class.

math.KT