SearcharxivSearch

arXiv subjects

Maissam Barkeshli

Publications and source records attributed to Maissam Barkeshli.

At least 19 recordsLinked to original sources

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate

Hyperparameter transfer allows extrapolating optimal optimization hyperparameters from small to large scales, making it critical for training large language models (LLMs). This is done either by fitting a scaling law to the hyperparameters or by a judicious choice of parameterization, such as Maximal Update ($μ$P), that renders optimal hyperparameters approximately scale invariant. In this paper, we first develop a framework to quantify hyperparameter transfer through three metrics: (1) the quality of the scaling law fit, (2) the robustness to extrapolation errors, and (3) the asymptotic loss penalty due to choice of parameterization. Next, we investigate through a comprehensive series of ablations why $μ$P appears to offer high-quality learning rate transfer relative to standard parameterization (SP), as existing theory is inadequate. We find that the overwhelming benefit of $μ$P relative to SP when training with AdamW arises simply from maximizing the learning rate of the embedding layer. In SP, the embedding layer learning rate acts as a bottleneck that induces training instabilities; increasing it by a factor of width to match $μ$P dramatically smooths out training while improving hyperparameter transfer. We also find that weight decay improves the scaling law fits, while, in the fixed token-per-parameter setting, it hurts the robustness of the extrapolation.

cs.LG

Crystalline topological invariants in quantum many-body systems

Crystalline symmetries give rise to topological invariants that can distinguish quantum phases of matter. Understanding these in strongly interacting systems is an ongoing research direction requiring non-perturbative methods. Recent developments have demonstrated that even classic models, like the Harper-Hofstadter model of free fermions on a lattice in a magnetic field, yield a host of crystalline symmetry protected topological invariants. Here we review some of these developments, focusing mainly on how to characterize, classify, and detect invariants arising from lattice translation and rotation symmetries along with charge conservation in two-dimensional systems, including integer and fractional Chern insulators.

cond-mat.str-el

Artificial Intelligence and the Structure of Mathematics

Recent progress in artificial intelligence (AI) is unlocking transformative capabilities for mathematics. There is great hope that AI will help solve major open problems and autonomously discover new mathematical concepts. In this essay, we further consider how AI may open a grand perspective on mathematics by forging a new route, complementary to mathematical\textbf{ logic,} to understanding the global structure of formal \textbf{proof}\textbf{s}. We begin by providing a sketch of the formal structure of mathematics in terms of universal proof and structural hypergraphs and discuss questions this raises about the foundational structure of mathematics. We then outline the main ingredients and provide a set of criteria to be satisfied for AI models capable of automated mathematical discovery. As we send AI agents to traverse Platonic mathematical worlds, we expect they will teach us about the nature of mathematics: both as a whole, and the small ribbons conducive to human understanding. Perhaps they will shed light on the old question: "Is mathematics discovered or invented?" Can we grok the terrain of these \textbf{Platonic worlds}?

cs.AI

Learning Pseudorandom Numbers with Transformers: Permuted Congruential Generators, Curricula, and Interpretability

We study the ability of Transformer models to learn sequences generated by Permuted Congruential Generators (PCGs), a widely used family of pseudo-random number generators (PRNGs). PCGs introduce substantial additional difficulty over linear congruential generators (LCGs) by applying a series of bit-wise shifts, XORs, rotations and truncations to the hidden state. We show that Transformers can nevertheless successfully perform in-context prediction on unseen sequences from diverse PCG variants, in tasks that are beyond published classical attacks. In our experiments we scale moduli up to $2^{22}$ using up to $50$ million model parameters and datasets with up to $5$ billion tokens. Surprisingly, we find even when the output is truncated to a single bit, it can be reliably predicted by the model. When multiple distinct PRNGs are presented together during training, the model can jointly learn them, identifying structures from different permutations. We demonstrate a scaling law with modulus $m$: the number of in-context sequence elements required for near-perfect prediction grows as $\sqrt{m}$. For larger moduli, optimization enters extended stagnation phases; in our experiments, learning moduli $m \geq 2^{20}$ requires incorporating training data from smaller moduli, demonstrating a critical necessity for curriculum learning. Finally, we analyze embedding layers and uncover a novel clustering phenomenon: the top principal components spontaneously group the integer inputs into bitwise rotationally-invariant clusters, revealing how representations can transfer from smaller to larger moduli.

cs.LG

Soft symmetries of topological orders

(2+1)D topological orders possess emergent symmetries given by a group $\text{Aut}(\mathcal{C})$, which consists of the braided tensor autoequivalences of the modular tensor category $\mathcal{C}$ that describes the anyons. In this paper we discuss cases where $\text{Aut}(\mathcal{C})$ has elements that neither permute anyons nor are associated to any symmetry fractionalization but are still non-trivial, which we refer to as soft symmetries. We point out that one can construct topological defects corresponding to such exotic symmetry actions by decorating with a certain class of gauged SPT states that cannot be distinguished by their torus partition function. This gives a physical interpretation to work by Davydov on soft braided tensor autoequivalences. This has a number of important implications for the classification of gapped boundaries, non-invertible spontaneous symmetry breaking, and the general classification of symmetry-enriched topological phases of matter. We also demonstrate analogous phenomena in higher dimensions, such as (3+1)D gauge theory with gauge group given by the quaternion group $Q_8$.

cond-mat.str-el

On the origin of neural scaling laws: from random graphs to natural language

Scaling laws have played a major role in the modern AI revolution, providing practitioners predictive power over how the model performance will improve with increasing data, compute, and number of model parameters. This has spurred an intense interest in the origin of neural scaling laws, with a common suggestion being that they arise from power law structure already present in the data. In this paper we study scaling laws for transformers trained to predict random walks (bigrams) on graphs with tunable complexity. We demonstrate that this simplified setting already gives rise to neural scaling laws even in the absence of power law structure in the data correlations. We further consider dialing down the complexity of natural language systematically, by training on sequences sampled from increasingly simplified generative language models, from 4,2,1-layer transformer language models down to language bigrams, revealing a monotonic evolution of the scaling exponents. Our results also include scaling laws obtained from training on random walks on random graphs drawn from Erdös-Renyi and scale-free Barabási-Albert ensembles. Finally, we revisit conventional scaling laws for language modeling, demonstrating that several essential results can be reproduced using 2 layer transformers with context length of 50, provide a critical analysis of various fits used in prior literature, demonstrate an alternative method for obtaining compute optimal curves as compared with current practice in published literature, and provide preliminary evidence that maximal update parameterization may be more parameter efficient than standard parameterization.

cs.LG

Invariants for (2+1)D bosonic crystalline topological insulators for all 17 wallpaper groups

We study bosonic symmetry-protected topological (SPT) phases in (2+1) dimensions with symmetry $G = G_{\text{space}}\times K$, where $G_{\text{space}}$ is a general wallpaper group and $K=\text{U}(1),\mathbb{Z}_N, \text{SO}(3)$ is an internal symmetry. In each case we propose a set of many-body invariants that can detect all the different phases predicted from real space constructions and group cohomology classifications. They are obtained by applying partial rotations and reflections to a given ground state, combined with suitable operations in $K$. The reflection symmetry invariants that we introduce include `double partial reflections', `weak partial reflections' and their `relative' or `twisted' versions which also depend on $K$. We verify our proposal through exact calculations on ground states constructed using real space constructions. We demonstrate our method in detail for the groups p4m and p4g, and in the case of p4m also derive a topological effective action involving gauge fields for orientation-reversing symmetries. Our results provide a concrete method to fully characterize (2+1)D crystalline topological invariants in bosonic SPT ground states.

cond-mat.str-el

(How) Can Transformers Predict Pseudo-Random Numbers?

Transformers excel at discovering patterns in sequential data, yet their fundamental limitations and learning mechanisms remain crucial topics of investigation. In this paper, we study the ability of Transformers to learn pseudo-random number sequences from linear congruential generators (LCGs), defined by the recurrence relation $x_{t+1} = a x_t + c \;\mathrm{mod}\; m$. We find that with sufficient architectural capacity and training data variety, Transformers can perform in-context prediction of LCG sequences with unseen moduli ($m$) and parameters ($a,c$). By analyzing the embedding layers and attention patterns, we uncover how Transformers develop algorithmic structures to learn these sequences in two scenarios of increasing complexity. First, we investigate how Transformers learn LCG sequences with unseen ($a, c$) but fixed modulus; and demonstrate successful learning up to $m = 2^{32}$. We find that models learn to factorize $m$ and utilize digit-wise number representations to make sequential predictions. In the second, more challenging scenario of unseen moduli, we show that Transformers can generalize to unseen moduli up to $m_{\text{test}} = 2^{16}$. In this case, the model employs a two-step strategy: first estimating the unknown modulus from the context, then utilizing prime factorizations to generate predictions. For this task, we observe a sharp transition in the accuracy at a critical depth $d= 3$. We also find that the number of in-context sequence elements needed to reach high accuracy scales sublinearly with the modulus.

cs.LG

Disclinations, dislocations, and emanant flux at Dirac criticality

What happens when fermions hop on a lattice with crystalline defects? The answer depends on topological quantum numbers which specify the action of lattice rotations and translations in the low energy theory. One can understand the topological quantum numbers as a twist of continuum gauge fields in terms of crystalline gauge fields. We find that disclinations and dislocations -- defects of crystalline symmetries -- generally lead in the continuum to a certain ``emanant'' quantized magnetic flux. To demonstrate these facts, we study in detail tight-binding models whose low-energy descriptions are (2+1)D Dirac cones. Our map from lattice to continuum defects explains the crystalline topological response to disclinations and dislocations, and motivates the fermion crystalline equivalence principle used in the classification of crystalline topological phases. When the gap closes, the presence of emanant flux leads to pair creation from the vacuum with the particles and anti-particles swirling around the defect. We compute the associated currents and energy density using the tools of defect conformal field theory. There is a rich set of renormalization group fixed points, depending on how particles scatter from the defect. At half flux, there is a defect conformal manifold leading to a continuum of possible low-energy theories. We present extensive numerical evidence supporting the emanant magnetic flux at lattice defects and we test our map between lattice and continuum defects in detail. We also point out a no-go result, which implies that a single (2+1)D Dirac cone in symmetry class AII is incompatible with a commuting $C_M$ rotational symmetry with $(C_M)^M = +1$.

cond-mat.str-el

When Can You Get Away with Low Memory Adam?

Adam is the go-to optimizer for training modern machine learning models, but it requires additional memory to maintain the moving averages of the gradients and their squares. While various low-memory optimizers have been proposed that sometimes match the performance of Adam, their lack of reliability has left Adam as the default choice. In this work, we apply a simple layer-wise Signal-to-Noise Ratio (SNR) analysis to quantify when second-moment tensors can be effectively replaced by their means across different dimensions. Our SNR analysis reveals how architecture, training hyperparameters, and dataset properties impact compressibility along Adam's trajectory, naturally leading to $\textit{SlimAdam}$, a memory-efficient Adam variant. $\textit{SlimAdam}$ compresses the second moments along dimensions with high SNR when feasible, and leaves when compression would be detrimental. Through experiments across a diverse set of architectures and training scenarios, we show that $\textit{SlimAdam}$ matches Adam's performance and stability while saving up to $98\%$ of total second moments. Code for $\textit{SlimAdam}$ is available at https://github.com/dayal-kalra/low-memory-adam.

cs.LG

Electric polarization in Chern insulators: Unifying many-body and single-particle approaches

Recently, it has been established that Chern insulators possess an intrinsic two-dimensional electric polarization, despite having gapless edge states and non-localizable Wannier orbitals. This polarization, $\vec{P}_{\text{o}}$, can be defined in a many-body setting from various physical quantities, including dislocation charges, boundary charge distributions, and linear momentum. Importantly, there is a dependence on a choice of real-space origin $\text{o}$ within the unit cell. In contrast, Coh and Vanderbilt extended the single-particle Berry phase definition of polarization to Chern insulators by choosing an arbitrary point in momentum space, $\vec{k}_0$. In this paper, we unify these two approaches and show that when the real-space origin $\text{o}$ and momentum-space point $\vec{k}_0$ are appropriately chosen in relation to each other, the Berry phase and many-body definitions of polarization are equal.

cond-mat.str-el

Universal Sharpness Dynamics in Neural Network Training: Fixed Point Analysis, Edge of Stability, and Route to Chaos

In gradient descent dynamics of neural networks, the top eigenvalue of the loss Hessian (sharpness) displays a variety of robust phenomena throughout training. This includes early time regimes where the sharpness may decrease during early periods of training (sharpness reduction), and later time behavior such as progressive sharpening and edge of stability. We demonstrate that a simple $2$-layer linear network (UV model) trained on a single training example exhibits all of the essential sharpness phenomenology observed in real-world scenarios. By analyzing the structure of dynamical fixed points in function space and the vector field of function updates, we uncover the underlying mechanisms behind these sharpness trends. Our analysis reveals (i) the mechanism behind early sharpness reduction and progressive sharpening, (ii) the required conditions for edge of stability, (iii) the crucial role of initialization and parameterization, and (iv) a period-doubling route to chaos on the edge of stability manifold as learning rate is increased. Finally, we demonstrate that various predictions from this simplified model generalize to real-world scenarios and discuss its limitations.

cs.LG

Fractionally Quantized Electric Polarization and Discrete Shift of Crystalline Fractional Chern Insulators

Fractional Chern insulators (FCI) with crystalline symmetry possess topological invariants that fundamentally have no analog in continuum fractional quantum Hall (FQH) states. Here we demonstrate through numerical calculations on model wave functions that FCIs possess a fractionally quantized electric polarization, $\vec{\mathscr{P}}_{\text{o}}$, where $\text{o}$ is a high symmetry point. $\vec{\mathscr{P}}_{\text{o}}$ takes fractional values as compared to the allowed values for integer Chern insulators because of the possibility that anyons carry fractional quantum numbers under lattice translation symmetries. $\vec{\mathscr{P}}_{\text{o}}$, together with the discrete shift $\mathscr{S}_{\text{o}}$, determine fractionally quantized universal contributions to electric charge in regions containing lattice disclinations, dislocations, boundaries, and/or corners, and which are fractions of the minimal anyon charge. We demonstrate how these invariants can be extracted using Monte Carlo computations on model wave functions with lattice defects for 1/2-Laughlin and 1/3-Laughlin FCIs on the square and honeycomb lattice, respectively, obtained using the parton construction. These results comprise a class of fractionally quantized response properties of topologically ordered states that go beyond the known ones discovered over thirty years ago.

cond-mat.str-el

Why Warmup the Learning Rate? Underlying Mechanisms and Improvements

It is common in deep learning to warm up the learning rate $η$, often by a linear schedule between $η_{\text{init}} = 0$ and a predetermined target $η_{\text{trgt}}$. In this paper, we show through systematic experiments using SGD and Adam that the overwhelming benefit of warmup arises from allowing the network to tolerate larger $η_{\text{trgt}}$ {by forcing the network to more well-conditioned areas of the loss landscape}. The ability to handle larger $η_{\text{trgt}}$ makes hyperparameter tuning more robust while improving the final performance. We uncover different regimes of operation during the warmup period, depending on whether training starts off in a progressive sharpening or sharpness reduction phase, which in turn depends on the initialization and parameterization. Using these insights, we show how $η_{\text{init}}$ can be properly chosen by utilizing the loss catapult mechanism, which saves on the number of warmup steps, in some cases completely eliminating the need for warmup. We also suggest an initialization for the variance in Adam which provides benefits similar to warmup.

cs.LG

Electric polarization and discrete shift from boundary and corner charge in crystalline Chern insulators

Recently, it has been shown how topological phases of matter with crystalline symmetry and $U(1)$ charge conservation can be partially characterized by a set of many-body invariants, the discrete shift $\mathscr{S}_{\text{o}}$ and electric polarization $\vec{\mathscr{P}}_{\text{o}}$, where $\text{o}$ labels a high symmetry point. Crucially, these can be defined even with non-zero Chern number and/or magnetic field. One manifestation of these invariants is through quantized fractional contributions to the charge in the vicinity of a lattice disclination or dislocation. In this paper, we show that these invariants can also be extracted from the length and corner dependence of the total charge (mod 1) on the boundary of the system. We provide a general formula in terms of $\mathscr{S}_{\text{o}}$ and $\vec{\mathscr{P}}_{\text{o}}$ for the total charge of any subregion of the system which can include full boundaries or bulk lattice defects, unifying boundary, corner, disclination, and dislocation charge responses into a single general theory. These results hold for Chern insulators, despite their gapless chiral edge modes, and for which an unambiguous definition of an intrinsically two-dimensional electric polarization has been unclear until recently. We also discuss how our theory can fully characterize the topological response of quadrupole insulators.

cond-mat.str-el

On stability of k-local quantum phases of matter

The current theoretical framework for topological phases of matter is based on the thermodynamic limit of a system with geometrically local interactions. A natural question is to what extent the notion of a phase of matter remains well-defined if we relax the constraint of geometric locality, and replace it with a weaker graph-theoretic notion of $k$-locality. As a step towards answering this question, we analyze the stability of the energy gap to perturbations for Hamiltonians corresponding to general quantum low-density parity-check codes, extending work of Bravyi and Hastings [Commun. Math. Phys. 307, 609 (2011)]. A corollary of our main result is that if there exist constants $\varepsilon_1,\varepsilon_2>0$ such that the size $Γ(r)$ of balls of radius $r$ on the interaction graph satisfy $Γ(r) = O(\exp(r^{1-\varepsilon_1}))$ and the local ground states of balls of radius $r\leρ^\ast = O(\log(n)^{1+\varepsilon_2})$ are locally indistinguishable, then the energy gap of the associated Hamiltonian is stable against local perturbations. This gives an almost exponential improvement over the $D$-dimensional Euclidean case, which requires $Γ(r) = O(r^D)$ and $ρ^\ast = O(n^α)$ for some $α> 0$. The approach we follow falls just short of proving stability of finite-rate qLDPC codes, which have $\varepsilon_1 = 0$; we discuss some strategies to extend the result to these cases. We discuss implications for the third law of thermodynamics, as $k$-local Hamiltonians can have extensive zero-temperature entropy.

quant-ph

Crystalline invariants of fractional Chern insulators

In the presence of crystalline symmetry, topologically ordered states can acquire a host of symmetry-protected invariants. These determine the patterns of crystalline symmetry fractionalization of the anyons in addition to fractionally quantized responses to lattice defects. Here we show how ground state expectation values of partial rotations centered at high symmetry points can be used to extract crystalline invariants. Using methods from conformal field theory and G-crossed braided tensor categories, we develop a theory of invariants obtained from partial rotations, which apply to both Abelian and non-Abelian topological orders. We then perform numerical Monte Carlo calculations for projected parton wave functions of fractional Chern insulators, demonstrating remarkable agreement between theory and numerics. For the topological orders we consider, we show that the Hall conductivity, filling fraction, and partial rotation invariants fully characterize the crystalline invariants of the system. Our results also yield invariants of continuum fractional quantum Hall states protected by spatial rotational symmetry.

cond-mat.str-el

Higher-group symmetry of (3+1)D fermionic $\mathbb{Z}_2$ gauge theory: logical CCZ, CS, and T gates from higher symmetry

It has recently been understood that the complete global symmetry of finite group topological gauge theories contains the structure of a higher-group. Here we study the higher-group structure in (3+1)D $\mathbb{Z}_2$ gauge theory with an emergent fermion, and point out that pumping chiral $p+ip$ topological states gives rise to a $\mathbb{Z}_{8}$ 0-form symmetry with mixed gravitational anomaly. This ordinary symmetry mixes with the other higher symmetries to form a 3-group structure, which we examine in detail. We then show that in the context of stabilizer quantum codes, one can obtain logical CCZ and CS gates by placing the code on a discretization of $T^3$ (3-torus) and $T^2 \rtimes_{C_2} S^1$ (2-torus bundle over the circle) respectively, and pumping $p+ip$ states. Our considerations also imply the possibility of a logical $T$ gate by placing the code on $\mathbb{RP}^3$ and pumping a $p+ip$ topological state.

cond-mat.str-el