SearcharxivSearch

arXiv subjects

Julian Gold

Publications and source records attributed to Julian Gold.

9 recordsLinked to original sources

Learning is Forgetting: LLM Training As Lossy Compression

Despite the increasing prevalence of large language models (LLMs), we still have a limited understanding of how their representational spaces are structured. This limits our ability to interpret how and what they learn or relate them to learning in humans. We argue LLMs are best seen as an instance of lossy compression, where over training they learn by retaining only information in their training data relevant to their objective(s). We show pre-training results in models that are optimally compressed for next-sequence prediction, approaching the Information Bottleneck bound on compression. Across an array of open weights models, each compresses differently, likely due to differences in the data and training recipes used. However even across different families of LLMs the optimality of a model's compression, and the information present in it, can predict downstream performance on across a wide array of benchmarks, letting us directly link representational structure to actionable insights about model performance. In the general case the work presented here offers a unified Information-Theoretic framing for how these models learn that is deployable at scale.

cs.LG

Hierarchical Refinement: Optimal Transport to Infinity and Beyond

Optimal transport (OT) has enjoyed great success in machine learning as a principled way to align datasets via a least-cost correspondence, driven in large part by the runtime efficiency of the Sinkhorn algorithm (Cuturi, 2013). However, Sinkhorn has quadratic space and time complexity in the number of points, limiting scalability to larger datasets. Low-rank OT achieves linear complexity, but by definition, cannot compute a one-to-one correspondence between points. When the optimal transport problem is an assignment problem between datasets then an optimal mapping, known as the Monge map, is guaranteed to be a bijection. In this setting, we show that the factors of an optimal low-rank coupling co-cluster each point with its image under the Monge map. We leverage this invariant to derive an algorithm, Hierarchical Refinement (HiRef), that dynamically constructs a multiscale partition of each dataset using low-rank OT subproblems, culminating in the bijective Monge map. Hierarchical Refinement runs in log-linear time and linear space, retaining the advantages of low-rank OT while overcoming its limited resolution. We demonstrate the advantages of Hierarchical Refinement on several datasets, including ones containing over a million points, scaling full-rank OT to problems previously beyond Sinkhorn's reach.

cs.LG

Low-Rank Optimal Transport through Factor Relaxation with Latent Coupling

Optimal transport (OT) is a general framework for finding a minimum-cost transport plan, or coupling, between probability distributions, and has many applications in machine learning. A key challenge in applying OT to massive datasets is the quadratic scaling of the coupling matrix with the size of the dataset. [Forrow et al. 2019] introduced a factored coupling for the k-Wasserstein barycenter problem, which [Scetbon et al. 2021] adapted to solve the primal low-rank OT problem. We derive an alternative parameterization of the low-rank problem based on the $\textit{latent coupling}$ (LC) factorization previously introduced by [Lin et al. 2021] generalizing [Forrow et al. 2019]. The LC factorization has multiple advantages for low-rank OT including decoupling the problem into three OT problems and greater flexibility and interpretability. We leverage these advantages to derive a new algorithm $\textit{Factor Relaxation with Latent Coupling}$ (FRLC), which uses $\textit{coordinate}$ mirror descent to compute the LC factorization. FRLC handles multiple OT objectives (Wasserstein, Gromov-Wasserstein, Fused Gromov-Wasserstein), and marginal constraints (balanced, unbalanced, and semi-relaxed) with linear space complexity. We provide theoretical results on FRLC, and demonstrate superior performance on diverse applications -- including graph clustering and spatial transcriptomics -- while demonstrating its interpretability.

cs.LG

On the number and size of holes in the growing ball of first-passage percolation

First-passage percolation is a random growth model defined on $\mathbb{Z}^d$ using i.i.d. nonnegative weights $(τ_e)$ on the edges. Letting $T(x,y)$ be the distance between vertices $x$ and $y$ induced by the weights, we study the random ball of radius $t$ centered at the origin, $B(t) = \{x \in \mathbb{Z}^d : T(0,x) \leq t\}$. It is known that for all such $τ_e$, the number of vertices (volume) of $B(t)$ is at least order $t^d$, and under mild conditions on $τ_e$, this volume grows like a deterministic constant times $t^d$. Defining a hole in $B(t)$ to be a bounded component of the complement $B(t)^c$, we prove that if $τ_e$ is not deterministic, then a.s., for all large $t$, $B(t)$ has at least $ct^{d-1}$ many holes, and the maximal volume of any hole is at least $c\log t$. Conditionally on the (unproved) uniform curvature assumption, we prove that a.s., for all large $t$, the number of holes is at most $(\log t)^C t^{d-1}$, and for $d=2$, no hole in $B(t)$ has volume larger than $(\log t)^C$. Without curvature, we show that no hole has volume larger than $Ct \log t$.

math.PR

The number of saddles of the spherical $p$-spin model

We show that the quenched complexity of saddles of the spherical pure $p$-spin model agrees with the annealed complexity when both are positive. Precisely, we show that the second moment of the number of critical values of a given finite index in a given interval has twice the growth rate of the first moment.

math.PR

Dynamical Freezing in a Spin Glass System with Logarithmic Correlations

We consider a continuous time random walk on the two-dimensional discrete torus, whose motion is governed by the discrete Gaussian free field on the corresponding box acting as a potential. More precisely, at any vertex the walk waits an exponentially distributed time with mean given by the exponential of the field and then jumps to one of its neighbors, chosen uniformly at random. We prove that throughout the low-temperature regime and at in-equilibrium timescales, the process admits a scaling limit as a spatial K-process driven by a random trapping landscape, which is explicitly related to the limiting extremal process of the field. Alternatively, the limiting process is a supercritical Liouville Brownian motion with respect to the continuum Gaussian free field on the box. This demonstrates rigorously and for the first time, as far as we know, a dynamical freezing in a spin glass system with logarithmically correlated energy levels.

math.PR

Isoperimetry in supercritical bond percolation in dimensions three and higher

We study the isoperimetric subgraphs of the infinite cluster $\textbf{C}_\infty$ for supercritical bond percolation on $\mathbb{Z}^d$ with $d\geq 3$. Specifically, we consider the subgraphs of $\textbf{C}_\infty \cap [-n,n]^d$ which have minimal open edge boundary to volume ratio. We prove a shape theorem for these subgraphs, obtaining that when suitably rescaled, these subgraphs converge almost surely to a translate of a deterministic shape. This deterministic shape is itself an isoperimetric set for a norm we construct. As a corollary, we obtain sharp asymptotics on a natural modification of the Cheeger constant for $\textbf{C}_\infty \cap [-n,n]^d$. This settles a conjecture of Benjamini for the version of the Cheeger constant defined here.

math.PR

Intrinsic isoperimetry of the giant component of supercritical bond percolation in dimension two

We study the isoperimetric subgraphs of the giant component $\textbf{C}_n$ of supercritical bond percolation on the square lattice. These are subgraphs of $\textbf{C}_n$ having minimal edge boundary to volume ratio. In contrast to the work of Biskup, Louidor, Procaccia and Rosenthal, the edge boundary is taken only within $\textbf{C}_n$ instead of the full infinite cluster. The isoperimetric subgraphs are shown to converge almost surely, after rescaling, to the collection of optimizers of a continuum isoperimetric problem emerging naturally from the model. We also show that the Cheeger constant of $\textbf{C}_n$ scales to a deterministic constant, which is itself an isoperimetric ratio, settling a conjecture of Benjamini in dimension two.

math.PR

A bound for orderings of Reidemeister moves

We provide an upper bound on the number of ordered Reidemeister moves required to pass between two diagrams of the same link. This bound is in terms of the number of unordered Reidemeister moves required.

math.GT