SearcharxivSearch

arXiv subjects

Itay Lavie

Publications and source records attributed to Itay Lavie.

5 recordsLinked to original sources

Statistical Properties of Training & Generalization

Deep learning has managed to evade numerous intuitions from classical statistics to achieve unprecedented performance on a number of real-world tasks. In this article, we investigate the key features and surprises of deep learning from a physics-informed perspective, taking care to point out and justify where possible the many choices inherent in constructing a deep learning model. In particular, we review the phenomenon of neural scaling laws and discuss their interplay with the constraints and inductive biases which may be present when applying machine learning to problems in physics.

stat.ML

Phase Transitions in Attention: A Bayesian Theory of Copy Head Emergence

Attention is the key mechanism underlying in-context learning in transformers, and attention patterns have been observed empirically to emerge abruptly during training. We present a Bayesian theory of feature learning in attention; we then focus on how the copy subcircuit in the first layer of an induction head is learned by analyzing a single-layer softmax attention network trained on a copy task. We derive a closed-form posterior over the attention matrix and reduce it to a low-dimensional order parameter space. This reduction reveals a phase transition in the amount of training data, which we verify using both Bayesian sampling and standard training with Adam. We contrast our results with linear attention and find that softmax attention exhibits a \emph{first-order phase transition} while in linear attention an initial \emph{second-order phase transition} is followed by a smooth, continuous evolution toward the structured attention pattern (\emph{crossover}). Our work provides a first-principles theoretical account of the abrupt emergence of the copy subcircuit, reminiscent of the one observed in training large language models.

stat.ML

Demystifying Spectral Bias on Real-World Data

Kernel ridge regression (KRR) and Gaussian processes (GPs) are fundamental tools in statistics and machine learning, with recent applications to highly over-parameterized deep neural networks. The ability of these tools to learn a target function is directly related to the eigenvalues of their kernel sampled on the input data distribution. Targets that have support on higher eigenvalues are more learnable. However, solving such eigenvalue problems on real-world data remains a challenge. Here, we consider cross-dataset learnability and show that one may use eigenvalues and eigenfunctions associated with highly idealized data measures to reveal spectral bias on complex datasets and bound learnability on real-world data. This allows us to leverage various symmetries that realistic kernels manifest to unravel their spectral bias.

stat.ML

Towards Understanding Inductive Bias in Transformers: A View From Infinity

We study inductive bias in Transformers in the infinitely over-parameterized Gaussian process limit and argue transformers tend to be biased towards more permutation symmetric functions in sequence space. We show that the representation theory of the symmetric group can be used to give quantitative analytical predictions when the dataset is symmetric to permutations between tokens. We present a simplified transformer block and solve the model at the limit, including accurate predictions for the learning curves and network outputs. We show that in common setups, one can derive tight bounds in the form of a scaling law for the learnability as a function of the context length. Finally, we argue WikiText dataset, does indeed possess a degree of permutation symmetry.

cs.LG

Roadmap to Thermal Dark Matter Beyond the WIMP Unitarity Bound

We study the general properties of the freezeout of a thermal relic. We give analytic estimates of the relic abundance for an arbitrary freezeout process, showing when instantaneous freezeout is appropriate and how it can be corrected when freezeout is slow. This is used to generalize the relationship between the dark mater mass and coupling that matches the observed abundance. The result encompasses well-studied particular examples, such as WIMPs, SIMPs, coannihilation, coscattering, inverse decays, and forbidden channels, and generalizes beyond them. In turn, this gives an approximate perturbative unitarity bound on the dark matter mass for an arbitrary thermal freezeout process. We show that going beyond the maximal masses allowed for freezeout via dark matter self-annihilations (WIMP-like, $m_{\rm DM}\gg\mathcal{O}(100~\rm TeV)$) predicts that there are nearly degenerate states with the dark matter and that the dark matter is generically metastable. We show how freezeout of a thermal relic may allow for dark matter masses up to the Planck scale.

hep-ph