SearcharxivSearch

arXiv subjects

Paolo Glorioso

Publications and source records attributed to Paolo Glorioso.

At least 19 recordsLinked to original sources

Diffusion cascade in a model of interacting random walkers

We consider the relaxation of finite-wavevector density waves in a facilitated classical lattice gas. Linear hydrodynamics predicts that such perturbations should relax exponentially, but nonlinear effects were predicted to cause subexponential relaxation via nonperturbative long-time tails. We present a detailed numerical study of this effect. While our results clearly indicate the importance of nonlinear effects, we find that the wavevector-dependence of the late-time relaxation is clearly inconsistent with theoretical predictions. We discuss manifestations of hydrodynamic nonlinearities in mesoscopic samples and at short times.

cond-mat.stat-mech

The Unreasonable Ineffectiveness of the Deeper Layers

How is knowledge stored in an LLM's weights? We study this via layer pruning: if removing a certain layer does not affect model performance in common question-answering benchmarks, then the weights in that layer are not necessary for storing the knowledge needed to answer those questions. To find these unnecessary parameters, we identify the optimal block of layers to prune by considering similarity across layers; then, to "heal" the damage, we perform a small amount of finetuning. Surprisingly, with this method we find minimal degradation of performance until after a large fraction (up to half) of the layers are removed for some common open-weight models. From a scientific perspective, the robustness of these LLMs to the deletion of layers implies either that current pretraining methods are not properly leveraging the parameters in the deeper layers of the network or that the shallow layers play a critical role in storing knowledge. For our study, we use parameter-efficient finetuning (PEFT) methods, specifically quantization and Low Rank Adapters (QLoRA), such that each of our experiments can be performed on a single 40GB A100 GPU.

cs.CL

The Zamba2 Suite: Technical Report

In this technical report, we present the Zamba2 series -- a suite of 1.2B, 2.7B, and 7.4B parameter hybrid Mamba2-transformer models that achieve state of the art performance against the leading open-weights models of their class, while achieving substantial gains in inference latency, throughput, and memory efficiency. The Zamba2 series builds upon our initial work with Zamba1-7B, optimizing its architecture, training and annealing datasets, and training for up to three trillion tokens. We provide open-source weights for all models of the Zamba2 series as well as instruction-tuned variants that are strongly competitive against comparable instruct-tuned models of their class. We additionally open-source the pretraining dataset, which we call Zyda-2, used to train the Zamba2 series of models. The models and datasets used in this work are openly available at https://huggingface.co/Zyphra

cs.LG

Zyda-2: a 5 Trillion Token High-Quality Dataset

In this technical report, we present Zyda-2: a five trillion token dataset for language model pretraining. Zyda-2 was used to train our Zamba2 series of models which are state-of-the-art for their weight class. We build Zyda-2 by collating high-quality open-source tokens such as FineWeb and DCLM, then distilling them to the highest-quality subset via cross-deduplication and model-based quality filtering. Zyda-2 is released under a permissive open license, and is available at https://huggingface.co/datasets/Zyphra/Zyda-2

cs.CL

Zyda: A 1.3T Dataset for Open Language Modeling

The size of large language models (LLMs) has scaled dramatically in recent years and their computational and data requirements have surged correspondingly. State-of-the-art language models, even at relatively smaller sizes, typically require training on at least a trillion tokens. This rapid advancement has eclipsed the growth of open-source datasets available for large-scale LLM pretraining. In this paper, we introduce Zyda (Zyphra Dataset), a dataset under a permissive license comprising 1.3 trillion tokens, assembled by integrating several major respected open-source datasets into a single, high-quality corpus. We apply rigorous filtering and deduplication processes, both within and across datasets, to maintain and enhance the quality derived from the original datasets. Our evaluations show that Zyda not only competes favorably with other open datasets like Dolma, FineWeb, and RefinedWeb, but also substantially improves the performance of comparable models from the Pythia suite. Our rigorous data processing methods significantly enhance Zyda's effectiveness, outperforming even the best of its constituent datasets when used independently.

cs.CL

Zamba: A Compact 7B SSM Hybrid Model

In this technical report, we present Zamba, a novel 7B SSM-transformer hybrid model which achieves competitive performance against leading open-weight models at a comparable scale. Zamba is trained on 1T tokens from openly available datasets and is the best non-transformer model at this scale. Zamba pioneers a unique architecture combining a Mamba backbone with a single shared attention module, thus obtaining the benefits of attention at minimal parameter cost. Due to its architecture, Zamba is significantly faster at inference than comparable transformer models and requires substantially less memory for generation of long sequences. Zamba is pretrained in two phases: the first phase is based on existing web datasets, while the second one consists of annealing the model over high-quality instruct and synthetic datasets, and is characterized by a rapid learning rate decay. We open-source the weights and all checkpoints for Zamba, through both phase 1 and annealing phases.

cs.LG

BlackMamba: Mixture of Experts for State-Space Models

State-space models (SSMs) have recently demonstrated competitive performance to transformers at large-scale language modeling benchmarks while achieving linear time and memory complexity as a function of sequence length. Mamba, a recently released SSM model, shows impressive performance in both language modeling and long sequence processing tasks. Simultaneously, mixture-of-expert (MoE) models have shown remarkable performance while significantly reducing the compute and latency costs of inference at the expense of a larger memory footprint. In this paper, we present BlackMamba, a novel architecture that combines the Mamba SSM with MoE to obtain the benefits of both. We demonstrate that BlackMamba performs competitively against both Mamba and transformer baselines, and outperforms in inference and training FLOPs. We fully train and open-source 340M/1.5B and 630M/2.8B BlackMamba models on 300B tokens of a custom dataset. We show that BlackMamba inherits and combines both of the benefits of SSM and MoE architectures, combining linear-complexity generation from SSM with cheap and fast inference from MoE. We release all weights, checkpoints, and inference code open-source. Inference code at: https://github.com/Zyphra/BlackMamba

cs.CL

Space-time generalization of mutual information

The mutual information characterizes correlations between spatially separated regions of a system. Yet, in experiments we often measure dynamical correlations, which involve probing operators that are also separated in time. Here, we introduce a space-time generalization of mutual information which, by construction, satisfies several natural properties of the mutual information and at the same time characterizes correlations across subsystems that are separated in time. In particular, this quantity, that we call the \emph{space-time mutual information}, bounds all dynamical correlations. We construct this quantity based on the idea of the quantum hypothesis testing. As a by-product, our definition provides a transparent interpretation in terms of an experimentally accessible setup. We draw connections with other notions in quantum information theory, such as quantum channel discrimination. Finally, we study the behavior of the space-time mutual information in several settings and contrast its long-time behavior in many-body localizing and thermalizing systems.

quant-ph

Generalized time-reversal symmetry and effective theories for nonequilibrium matter

The past decade has witnessed the development of systematic effective theories for dissipative thermal systems. Here, we describe an analogous effective theory framework that applies to the classical stochastic dynamics of nonequilibrium systems. We illustrate this approach using a range of examples, including nonreciprocal (predator-prey) dynamics, dissipative and driven rigid-body motion, and active chiral fluids and solids. Many of these systems exhibit a generalized time-reversal symmetry, which plays a crucial role within our formalism, and in many cases can be implemented within the Martin-Siggia-Rose path integral. This effective theory formalism yields generalizations of the fluctuation-dissipation theorem and second law of thermodynamics valid out of equilibrium. By stipulating a stationary distribution and a set of symmetries -- rather than postulating the stochastic equations of motion directly -- this formalism provides an alternative route to building phenomenological models of driven and active matter. We hope that this approach facilitates a systematic investigation of the universality classes of active matter, and provides a common language for nonequilibrium many-body physics from high energy to condensed matter.

cond-mat.stat-mech

Flatter, faster: scaling momentum for optimal speedup of SGD

Commonly used optimization algorithms often show a trade-off between good generalization and fast training times. For instance, stochastic gradient descent (SGD) tends to have good generalization; however, adaptive gradient methods have superior training times. Momentum can help accelerate training with SGD, but so far there has been no principled way to select the momentum hyperparameter. Here we study training dynamics arising from the interplay between SGD with label noise and momentum in the training of overparametrized neural networks. We find that scaling the momentum hyperparameter $1-β$ with the learning rate to the power of $2/3$ maximally accelerates training, without sacrificing generalization. To analytically derive this result we develop an architecture-independent framework, where the main assumption is the existence of a degenerate manifold of global minimizers, as is natural in overparametrized models. Training dynamics display the emergence of two characteristic timescales that are well-separated for generic values of the hyperparameters. The maximum acceleration of training is reached when these two timescales meet, which in turn determines the scaling limit we propose. We confirm our scaling rule for synthetic regression problems (matrix sensing and teacher-student paradigm) and classification for realistic datasets (ResNet-18 on CIFAR10, 6-layer MLP on FashionMNIST), suggesting the robustness of our scaling rule to variations in architectures and datasets.

cs.LG

Cross Entropy Benchmark for Measurement-Induced Phase Transitions

We investigate the prospects of employing the linear cross-entropy to experimentally access measurement-induced phase transitions (MIPT) without requiring any postselection of quantum trajectories. For two random circuits that are identical in the bulk but with different initial states, the linear cross-entropy $χ$ between the bulk measurement outcome distributions in the two circuits acts as a boundary order parameter, and can be used to distinguish the volume law from area law phases. In the volume law phase (and in the thermodynamic limit) the bulk measurements cannot distinguish between the two different initial states, and $χ= 1$. In the area law phase $χ< 1$. For circuits with Clifford gates, we provide numerical evidence that $χ$ can be sampled to accuracy $ε$ from $O(1/ε^2)$ trajectories, by running the first circuit on a quantum simulator without postselection, aided by a classical simulation of the second. We also find that for weak depolarizing noise the signature of the MIPT is still present for intermediate system sizes. In our protocol we have the freedom of choosing initial states such that the "classical" side can be simulated efficiently, while simulating the "quantum" side is still classically hard.

quant-ph

Kardar-Parisi-Zhang Universality at the Edge of Laughlin States

In this letter, we investigate the dissipative dynamics at the edge of Laughlin fractional quantum Hall (FQH) states starting from the hydrodynamic framework of the composite Boson theory recently developed in arXiv:2203.06516. Critical to this description is the choice of boundary conditions, which ultimately stems from the choice of hydrodynamic variables in terms of condensate degrees of freedom. Given the gapped nature of bulk, one would expect dissipation effects to play an important role only near the FQH edge. Thus, one envisions a scenario where the bulk hydro equations remain unmodified, while the dissipation effects are introduced at the edge via boundary conditions. We have recently shown that the anomaly requirements fix the boundary conditions of the FQH fluid to be no-penetration and no-stress boundary conditions. In this work, we introduce energy dissipation in the no-stress boundary condition leading to charge diffusion at the boundary. The resulting dissipative edge dynamics is quite rigid from a hydro perspective, as it has to preserve the edge charge continuity and the anomaly structure. We show that the diffusive edge dynamics with fluctuation-dissipation relations within a power counting scheme belong to the Kardar-Parisi-Zhang universality class.

cond-mat.mes-hall

Goldstone bosons and fluctuating hydrodynamics with dipole and momentum conservation

We develop a Schwinger-Keldysh effective field theory describing the hydrodynamics of a fluid with conserved charge and dipole moments, together with conserved momentum. The resulting hydrodynamic modes are highly unusual, including sound waves with quadratic (magnon-like) dispersion relation and subdiffusive decay rate. Hydrodynamics itself is unstable below four spatial dimensions. We show that the momentum density is, at leading order, the Goldstone boson for a dipole symmetry which appears spontaneously broken at finite charge density. Unlike an ordinary fluid, the presence or absence of energy conservation qualitatively changes the decay rates of the hydrodynamic modes. This effective field theory naturally couples to curved spacetime and background gauge fields; in the flat spacetime limit, we reproduce the "mixed rank tensor fields" previously coupled to fracton matter.

hep-th

Fracton superfluid hydrodynamics

We examine the hydrodynamics of systems with spontaneously broken multipolar symmetries using a systematic effective field theory. We focus on the simplest non-trivial setting: a system with charge and dipole symmetry, but without momentum conservation. When no symmetries are broken, our formalism reproduces the quartic subdiffusion ($ω\sim -i k^4$) characteristic of `fracton hydrodynamics' with conserved dipole moment. Our formalism also captures spontaneous breaking of charge and/or dipole symmetry. When charge symmetry is spontaneously broken, the hydrodynamic modes are quadratically propagating and quartically relaxing ($ω\sim \pm k^2 - ik^4$). When the dipole symmetry is spontaneously broken but the charge symmetry is preserved, then we find quadratically relaxing (diffusive) transverse modes, plus another mode which depending on parameters may be either purely diffusive ($ω\sim -i k^2$) or quadratically propagating and quadratically relaxing ($ω\sim \pm k^2 -i k^2$). Our work provides concrete predictions that may be tested in near-term cold atom experiments, and also lays out a general framework that may be applied to study systems with spontaneously broken multipolar symmetries.

cond-mat.stat-mech

Far-from-equilibrium universality in two-dimensional Heisenberg antiferromagnets

We study the far-from-equilibrium dynamics of isolated two-dimensional Heisenberg antiferromagnets. We consider spin spiral initial conditions which imprint a position-dependent staggered-magnetization (or Neel order) in the two-dimensional lattice. Remarkably, we find a long-lived prethermal regime characterized by self-similar behavior of staggered magnetization fluctuations, although the system has no long-range order at finite energy and the staggered magnetization does not couple with conserved charges. Exploiting the separation of length scales introduced by the initial condition, we derive a simplified analytical model that allow us to compute the spatial-temporal scaling exponents and power-law distribution of the staggered magnetization fluctuations analytically, and find excellent agreement with numerical simulations using phase space methods. The scaling exponents are insensitive to details of the initial condition, in particular, no fine-tuning of energy is required to trigger the self-similar scaling regime. Compared with recent results on far-from-equilibrium universality on the Heisenberg ferromagnet, we find quantitatively distinct spatial-temporal scaling exponents, therefore suggesting that the same model with ferromagnetic and antiferromagnetic initial conditions can host different universal regimes. Our predictions are relevant to ultra-cold atoms simulators of Heisenberg magnets and driven antiferromagnetic insulators.

cond-mat.stat-mech

Joule heating in bad and slow metals

Heat supplied to a metal is absorbed by the electrons and then transferred to the lattice. In conventional metals energy is released to the lattice by phonons emitted from the Lindhard continuum. However in a `bad' metal, with short mean free path, the low energy Lindhard continuum is destroyed. Furthermore in a `slow' metal, with Fermi velocity less than the sound velocity, particle-hole pairs are kinematically unable to emit phonons. To describe energy transfer to the lattice in these cases we obtain a general Kubo formula for the energy relaxation rate in terms of the electronic density spectral weight $\text{Im} \, G^R_{nn}(ω_k,k)$ evaluated on the phonon dispersion $ω_k$. We apply our Kubo formula to the high temperature Hubbard model, using recent data from quantum Monte Carlo and experiments in ultracold atoms to characterize $\text{Im} \, G^R_{nn}(ω_k,k)$. We furthermore use recent data from electron energy-loss spectroscopy to estimate the energy relaxation rate of the cuprate strange metal to a high energy optical phonon. As a second, distinct, application of our formalism we consider `slow' metals. These are defined to have Fermi velocity less than the sound velocity, so that particle-hole pairs are kinematically unable to emit phonons. We obtain an expression for the energy relaxation rate of a slow metal in terms of the optical conductivity.

cond-mat.str-el

Fracton hydrodynamics without time-reversal symmetry

We present an effective field theory for the nonlinear fluctuating hydrodynamics of a single conserved charge with or without time-reversal symmetry, based on the Martin-Siggia-Rose formalism. Applying this formalism to fluids with only charge and multipole conservation, and with broken time-reversal symmetry, we predict infinitely many new dynamical universality classes, including some with arbitrarily large upper critical dimensions. Using large scale simulations of classical Markov chains, we find numerical evidence for a breakdown of hydrodynamics in quadrupole-conserving models with broken time-reversal symmetry in one spatial dimension.

cond-mat.stat-mech

Breakdown of hydrodynamics below four dimensions in a fracton fluid

We present the nonlinear fluctuating hydrodynamics which governs the late time dynamics of a chaotic many-body system with simultaneous charge/mass, dipole/center of mass, and momentum conservation. This hydrodynamic effective theory is unstable below four spatial dimensions: dipole-conserving fluids at rest become unstable to fluctuations, and are governed not by hydrodynamics, but by a fractonic generalization of the Kardar-Parisi-Zhang universality class. We numerically simulate many-body classical dynamics in one-dimensional models with dipole and momentum conservation, and find evidence for a breakdown of hydrodynamics, along with a new universality class of undriven yet non-equilbrium dynamics.

cond-mat.str-el