Searcharxiv⌕ Search

arXiv subjects

Lachlan Ewen MacDonald

Publications and source records attributed to Lachlan Ewen MacDonald.

15 recordsLinked to original sources

FastGuide: Accelerating Reward Guidance for Diffusion Large Language Models

Gradient-based reward guidance provides a flexible way to use downstream reward models to control masked diffusion language models at inference time. However, its computational cost remains high as each decoding iteration incurs expensive diffusion model forward passes and reward model backpropagation steps. To address this, we introduce FastGuide, an adaptive hybrid of parallel and autoregressive decoding to accelerate reward guidance for diffusion language models. In analogy to parallel decoding, FastGuide amortizes the cost of reward model backpropagation by computing guidance once per decoding step and reusing it to generate multiple tokens. Within each decoding step, FastGuide makes diffusion forward passes autoregressive by unmasking tokens one at a time while efficiently recomputing token distributions after each unmasking by utilizing KV caching techniques and sparse recomputation of attention. Lastly, to adapt hybrid decoding to the model's confidence, FastGuide defers any token that the model is unconfident about under its recomputed distribution. Experiments on three reward benchmarks demonstrate that FastGuide is up to $4.4\times$ faster than sequential reward-guided decoding while retaining similar generation quality.

cs.CL↗

ANaLOG: Anisotropic Native-Latent Operator Guidance for Solving Inverse Problems

Native-latent guidance is a recent paradigm for solving inverse problems with latent diffusion models. It replaces repeated evaluations of the image-space forward model, each requiring a decoder pass, with efficient guidance computed using a learned latent-space surrogate. However, existing methods apply guidance uniformly across latent dimensions, ignoring that measurements are informative only along certain directions and that the reliability of model predictions varies across inputs and timesteps. We propose ANaLOG, a framework for efficient uncertainty-aware guidance with pretrained latent diffusion models. ANaLOG models uncertainty by learning an anisotropic, input- and time-dependent covariance that is integrated into the guidance mechanism to emphasize reliable directions and downweight uncertain ones. We theoretically analyze this framework in a linear model setting and prove that anisotropic, uncertainty-aware weighting is necessary for correct sampling, whereas isotropic guidance induces sampling errors. Experiments across five challenging inverse problems show that ANaLOG improves perceptual reconstruction quality over existing methods while preserving efficiency.

eess.IV↗

Dynamics of Gradient Descent with Large Step Size Near a Manifold of Flat Minima

An important quantity in the theory of gradient descent (GD) is the \emph{sharpness}, defined as the largest eigenvalue of the objective Hessian. Classical analyses typically require the step size to be uniformly smaller than twice the reciprocal of the sharpness, but this condition is frequently violated in the training of deep neural networks. Recent work bridges this gap in the setting of overparametrised least-squares with a \emph{single scalar output}, providing a normal form for large-step GD in a neighbourhood of an \emph{isolated} flat minimum and establishing three corresponding convergence results. In this paper, we extend this theory in two directions: (1) to overparametrised least-squares with \emph{vector-valued outputs} (including regression with arbitrarily many observations), and (2) to a neighbourhood of a \emph{manifold} of flat minima (which we show is essential for applications such as matrix factorisation). We generalise both the normal form and all three convergence theorems of \cite{macdonaldeos} to this broader setting, overcoming several technical challenges. We further show that our framework applies to deep matrix factorisation under mild assumptions, yielding several new structural results. In particular, we prove that the set of flat minima forms a fibre bundle over a product of spheres, and that the sharpness is Morse-Bott along this manifold.

cs.LG↗

Convergence Rates for Gradient Descent on the Edge of Stability in Overparametrised Least Squares

Classical optimisation theory guarantees monotonic objective decrease for gradient descent (GD) when employed in a small step size, or ``stable", regime. In contrast, gradient descent on neural networks is frequently performed in a large step size regime called the ``edge of stability", in which the objective decreases non-monotonically with an observed implicit bias towards flat minima. In this paper, we take a step toward quantifying this phenomenon by providing convergence rates for gradient descent with large learning rates in an overparametrised least squares setting. The key insight behind our analysis is that, as a consequence of overparametrisation, the set of global minimisers forms a Riemannian manifold $M$, which enables the decomposition of the GD dynamics into components parallel and orthogonal to $M$. The parallel component corresponds to Riemannian gradient descent on the objective sharpness, while the orthogonal component is a bifurcating dynamical system. This insight allows us to derive convergence rates in three regimes characterised by the learning rate size: (a) the subcritical regime, in which transient instability is overcome in finite time before linear convergence to a suboptimally flat global minimum; (b) the critical regime, in which instability persists for all time with a power-law convergence toward the optimally flat global minimum; and (c) the supercritical regime, in which instability persists for all time with linear convergence to an orbit of period two centred on the optimally flat global minimum.

cs.LG↗

Disintegration theorem for multifunctions, with applications to empirical Wasserstein distances and average-case statistical bounds

We prove a generalisation of the disintegration theorem to the setting of multifunctions between Polish probability spaces. Whereas the classical disintegration theorem guarantees the disintegration of a probability measure along the partition of the underlying space by the fibres of a measurable function, our theorem gives necessary and sufficient conditions for the measure to disintegrate along a cover of the underlying space defined by the fibres of a measurable multifunction. Building on this theorem, we introduce a new statistical notion: We declare a metric Polish probability space to be asymptotically disintegrable if $n$ i.i.d.-centred balls of decreasing radius carry a disintegration of the measure with probability tending to unity as $n\rightarrow\infty$. We give a number of both $1$-dimensional and higher-dimensional examples of asymptotically disintegrable spaces with associated quantitative rates, as well as a strong counterexample. Finally, we give two applications of the notion of asymptotic disintegrability. First, we prove that asymptotically disintegrable spaces admit an easy high-probability quantification of the law of large numbers in Wasserstein space, which in all dimensions either recovers or improves upon the best known rates in some regimes, and is never any worse than existing rates by more than a factor of 2 in the exponent of $n$, where $n$ is the number of sample points. Second, we prove that any asymptotically disintegrable space admits a high-probability bound on the error in approximating the expectation of any Lipschitz function by its empirical average over an i.i.d.\ sample. The bound is average-case in the sense that it depends only on the empirical average of the local Lipschitz constants of the function, rather than the global Lipschitz constant as obtained by Kantorovich-Rubinstein duality.

math.PR↗

Chern-Weil theory for Haefliger-singular foliations

We give a Chern-Weil map for the Gel'fand-Fuks characteristic classes of Haefliger-singular foliations, those foliations defined by smooth Haefliger structures with dense regular set. Our characteristic map constructs, out of singular geometric structures adapted to singularities, explicit forms representing characteristic classes in de Rham cohomology. The forms are functorial under foliation morphisms. We prove that the theory applies, up to homotopy, to general smooth Haefliger structures: subject only to obvious necessary dimension constraints, every smooth Haefliger structure is homotopic to a Haefliger-singular foliation, and any morphism of Haefliger structures is homotopic to a morphism of Haefliger-singular foliations. As an application, we provide a generalisation to the singular setting of the classical construction of forms representing the Godbillon-Vey invariant.

math.DG↗

Understanding the Learning Dynamics of LoRA: A Gradient Flow Perspective on Low-Rank Adaptation in Matrix Factorization

Despite the empirical success of Low-Rank Adaptation (LoRA) in fine-tuning pre-trained models, there is little theoretical understanding of how first-order methods with carefully crafted initialization adapt models to new tasks. In this work, we take the first step towards bridging this gap by theoretically analyzing the learning dynamics of LoRA for matrix factorization (MF) under gradient flow (GF), emphasizing the crucial role of initialization. For small initialization, we theoretically show that GF converges to a neighborhood of the optimal solution, with smaller initialization leading to lower final error. Our analysis shows that the final error is affected by the misalignment between the singular spaces of the pre-trained model and the target matrix, and reducing the initialization scale improves alignment. To address this misalignment, we propose a spectral initialization for LoRA in MF and theoretically prove that GF with small spectral initialization converges to the fine-tuning task with arbitrary precision. Numerical experiments from MF and image classification validate our findings.

cs.LG↗

D'OH: Decoder-Only Random Hypernetworks for Implicit Neural Representations

Deep implicit functions have been found to be an effective tool for efficiently encoding all manner of natural signals. Their attractiveness stems from their ability to compactly represent signals with little to no offline training data. Instead, they leverage the implicit bias of deep networks to decouple hidden redundancies within the signal. In this paper, we explore the hypothesis that additional compression can be achieved by leveraging redundancies that exist between layers. We propose to use a novel runtime decoder-only hypernetwork - that uses no offline training data - to better exploit cross-layer parameter redundancy. Previous applications of hypernetworks with deep implicit functions have employed feed-forward encoder/decoder frameworks that rely on large offline datasets that do not generalize beyond the signals they were trained on. We instead present a strategy for the optimization of runtime deep implicit functions for single-instance signals through a Decoder-Only randomly projected Hypernetwork (D'OH). By directly changing the latent code dimension, we provide a natural way to vary the memory footprint of neural representations without the costly need for neural architecture search on a space of alternative low-rate structures.

cs.LG↗

On skip connections and normalisation layers in deep optimisation

We introduce a general theoretical framework, designed for the study of gradient optimisation of deep neural networks, that encompasses ubiquitous architecture choices including batch normalisation, weight normalisation and skip connections. Our framework determines the curvature and regularity properties of multilayer loss landscapes in terms of their constituent layers, thereby elucidating the roles played by normalisation layers and skip connections in globalising these properties. We then demonstrate the utility of this framework in two respects. First, we give the only proof of which we are aware that a class of deep neural networks can be trained using gradient descent to global optima even when such optima only exist at infinity, as is the case for the cross-entropy cost. Second, we identify a novel causal mechanism by which skip connections accelerate training, which we verify predictively with ResNets on MNIST, CIFAR10, CIFAR100 and ImageNet.

cs.LG↗

On progressive sharpening, flat minima and generalisation

We present a new approach to understanding the relationship between loss curvature and input-output model behaviour in deep learning. Specifically, we use existing empirical analyses of the spectrum of deep network loss Hessians to ground an ansatz tying together the loss Hessian and the input-output Jacobian over training samples during the training of deep neural networks. We then prove a series of theoretical results which quantify the degree to which the input-output Jacobian of a model approximates its Lipschitz norm over a data distribution, and deduce a novel generalisation bound in terms of the empirical Jacobian. We use our ansatz, together with our theoretical results, to give a new account of the recently observed progressive sharpening phenomenon, as well as the generalisation properties of flat minima. Experimental evidence is provided to validate our claims.

cs.LG↗

Flow supervision for Deformable NeRF

In this paper we present a new method for deformable NeRF that can directly use optical flow as supervision. We overcome the major challenge with respect to the computationally inefficiency of enforcing the flow constraints to the backward deformation field, used by deformable NeRFs. Specifically, we show that inverting the backward deformation function is actually not needed for computing scene flows between frames. This insight dramatically simplifies the problem, as one is no longer constrained to deformation functions that can be analytically inverted. Instead, thanks to the weak assumptions required by our derivation based on the inverse function theorem, our approach can be extended to a broad class of commonly used backward deformation field. We present results on monocular novel view synthesis with rapid object motion, and demonstrate significant improvements over baselines without flow supervision.

cs.CV↗

Enabling equivariance for arbitrary Lie groups

Although provably robust to translational perturbations, convolutional neural networks (CNNs) are known to suffer from extreme performance degradation when presented at test time with more general geometric transformations of inputs. Recently, this limitation has motivated a shift in focus from CNNs to Capsule Networks (CapsNets). However, CapsNets suffer from admitting relatively few theoretical guarantees of invariance. We introduce a rigourous mathematical framework to permit invariance to any Lie group of warps, exclusively using convolutions (over Lie groups), without the need for capsules. Previous work on group convolutions has been hampered by strong assumptions about the group, which precludes the application of such techniques to common warps in computer vision such as affine and homographic. Our framework enables the implementation of group convolutions over any finite-dimensional Lie group. We empirically validate our approach on the benchmark affine-invariant classification task, where we achieve 30% improvement in accuracy against conventional CNNs while outperforming most CapsNets. As further illustration of the generality of our framework, we train a homography-convolutional model which achieves superior robustness on a homography-perturbed dataset, where CapsNet results degrade.

cs.CV↗

The Holonomy Groupoids of Singularly Foliated Bundles

We define a notion of connection in a fibre bundle that is compatible with a singular foliation of the base. Fibre bundles equipped with such connections are in plentiful supply, arising naturally for any Lie groupoid-equivariant bundle, and simultaneously generalising regularly foliated bundles in the sense of Kamber-Tondeur and singular foliations. We define hierarchies of diffeological holonomy groupoids associated to such bundles, which arise from the parallel transport of jet/germinal conservation laws. We show that the groupoids associated in this manner to trivial singularly foliated bundles are quotients of Androulidakis-Skandalis holonomy groupoids, which coincide with Androulidakis-Skandalis holonomy groupoids in the regular case. Finally we prove functoriality of all our constructions under appropriate morphisms.

math.DG↗

Hierarchies of holonomy groupoids for foliated bundles

We give a new construction of the holonomy groupoid of a regular foliation in terms of a partial connection on a diffeological principal bundle of germs of transverse parametrisations. We extend these ideas to construct a novel holonomy groupoid for any foliated bundle, which we prove sits at the top of a hierarchy of diffeological jet holonomy groupoids associated to the foliated bundle.

math.DG↗