SearcharxivSearch

arXiv subjects

Ro Jefferson

Publications and source records attributed to Ro Jefferson.

At least 19 recordsLinked to original sources

Deep neural networks as lattice gauge theories

We modify the NN/QFT duality [1] to incorporate the layerwise permutation symmetry of the network, resulting in a $(0\!+\!1)$-dimensional lattice gauge theory, in which each layer of $N$ neurons acts as an $N$-component lattice site, and the weight matrices play the role of gauge fields living on the links. In this framework, we compute the tree-level neuron-neuron propagator which describes the evolution of layer variance in the network, and develop the Feynman diagram machinery to compute interactions in the perturbative expansion in $1/N$. In particular, we obtain a recursive expression for all corrections to the exact propagator at $O(1)$, representing statistical fluctuations in the ensemble of networks, including infinitely-many loop diagrams mediating the interactions from previous layers. We also present a preliminary analysis of neuron scattering amplitudes that contribute order-by-order in $1/N$, which provides a field-theoretic framework for studying higher-point correlations, and by extension information propagation, in deep networks. We remark on some interesting directions for future work at the intersection of neural networks and quantum field theory.

hep-th

Towards Worst-Case Guarantees with Scale-Aware Interpretability

Neural networks organize information according to the hierarchical, multi-scale structure of natural data. Methods to interpret model internals should be similarly scale-aware, explicitly tracking how features compose across resolutions and guaranteeing bounds on the influence of fine-grained structure that is discarded as irrelevant noise. We posit that the renormalisation framework from physics can meet this need by offering technical tools that can overcome limitations of current methods. Moreover, relevant work from adjacent fields has now matured to a point where scattered research threads can be synthesized into practical, theory-informed tools. To combine these threads in an AI safety context, we propose a unifying research agenda -- \emph{scale-aware interpretability} -- to develop formal machinery and interpretability tools that have robustness and faithfulness properties supported by statistical physics.

hep-th

Algebraic perturbation theory: traversable wormholes and generalized entropy beyond subleading order

The crossed product has recently emerged as an important ingredient in describing algebras of observables for quantum field theory and gravity. We combine this with perturbation theory, and study perturbative crossed product algebras obtained from a unitary deformation of the original system. Motivated by the problem of black hole evaporation, we propose an abstract framework in which black hole information can be transferred to Hawking radiation by passing to a perturbative crossed product exhibiting a degree of non-locality in its modular structure. As both a concrete example and a toy model for evaporation, we analyze the algebra of observables of the traversable wormhole in anti-de-Sitter space. We obtain new contributions to the generalized entropy beyond subleading order relative to the original work by Gao, Jafferis, and Wall. We close with some comments on the potential applicability of the algebraic approach to quantum gravity.

hep-th

Predictive Coding Networks and Inference Learning: Tutorial and Survey

Recent years have witnessed a growing call for renewed emphasis on neuroscience-inspired approaches in artificial intelligence research, under the banner of NeuroAI. A prime example of this is predictive coding networks (PCNs), based on the neuroscientific framework of predictive coding. This framework views the brain as a hierarchical Bayesian inference model that minimizes prediction errors through feedback connections. Unlike traditional neural networks trained with backpropagation (BP), PCNs utilize inference learning (IL), a more biologically plausible algorithm that explains patterns of neural activity that BP cannot. Historically, IL has been more computationally intensive, but recent advancements have demonstrated that it can achieve higher efficiency than BP with sufficient parallelization. Furthermore, PCNs can be mathematically considered a superset of traditional feedforward neural networks (FNNs), significantly extending the range of trainable architectures. As inherently probabilistic (graphical) latent variable models, PCNs provide a versatile framework for both supervised learning and unsupervised (generative) modeling that goes beyond traditional artificial neural networks. This work provides a comprehensive review and detailed formal specification of PCNs, particularly situating them within the context of modern ML methods. This positions PC as a promising framework for future ML innovations.

cs.LG

Opening the Black Box: predicting the trainability of deep neural networks with reconstruction entropy

An important challenge in machine learning is to predict the initial conditions under which a given neural network will be trainable. We present a method for predicting the trainable regime in parameter space for deep feedforward neural networks (DNNs) based on reconstructing the input from subsequent activation layers via a cascade of single-layer auxiliary networks. We show that a single epoch of training of the shallow cascade networks is sufficient to predict the trainability of the deep feedforward network on a range of datasets (MNIST, CIFAR10, FashionMNIST, and white noise), thereby providing a significant reduction in overall training time. We achieve this by computing the relative entropy between reconstructed images and the original inputs, and show that this probe of information loss is sensitive to the phase behaviour of the network. We further demonstrate that this method generalizes to residual neural networks (ResNets) and convolutional neural networks (CNNs). Moreover, our method illustrates the network's decision making process by displaying the changes performed on the input data at each layer, which we demonstrate for both a DNN trained on MNIST and the vgg16 CNN trained on the ImageNet dataset. Our results provide a technique for significantly accelerating the training of large neural networks.

cs.LG

Wilsonian Renormalization of Neural Network Gaussian Processes

Separating relevant and irrelevant information is key to any modeling process or scientific inquiry. Theoretical physics offers a powerful tool for achieving this in the form of the renormalization group (RG). Here we demonstrate a practical approach to performing Wilsonian RG in the context of Gaussian Process (GP) Regression. We systematically integrate out the unlearnable modes of the GP kernel, thereby obtaining an RG flow of the GP in which the data sets the IR scale. In simple cases, this results in a universal flow of the ridge parameter, which becomes input-dependent in the richer scenario in which non-Gaussianities are included. In addition to being analytically tractable, this approach goes beyond structural analogies between RG and neural networks by providing a natural connection between RG flow and learnable vs. unlearnable modes. Studying such flows may improve our understanding of feature learning in deep neural networks, and enable us to identify potential universality classes in these models.

cs.LG

Crossed product algebras and generalized entropy for subregions

An early result of algebraic quantum field theory is that the algebra of any subregion in a QFT is a von Neumann factor of type III$_1$, in which entropy cannot be well-defined because such algebras do not admit a trace or density states. However, associated to the algebra is a modular group of automorphisms characterizing the local dynamics of degrees of freedom in the region, and the crossed product of the algebra with its modular group yields a type II$_\infty$ factor, in which traces and hence von Neumann entropy can be well-defined. In this work, we generalize recent constructions of the crossed product algebra for the TFD to, in principle, arbitrary spacetime regions in arbitrary QFTs, paving the way to the study of entanglement entropy without UV divergences. In contrast to previous works, we emphasize that this construction is independent of gravity. In this sense, the crossed product construction represents a refinement of Haag's assignment of nets of observable algebras to spacetime regions by providing a natural construction of a type II factor. We present several concrete examples: a QFT in Rindler space, a CFT in an open ball of Minkowski space, and arbitrary boundary subregions in AdS/CFT. In the holographic setting, we provide a novel argument for why the bulk dual must be the entanglement wedge, and discuss the distinction arising from boundary modular flow between causal and entanglement wedges for excited states and disjoint regions.

hep-th

Criticality versus uniformity in deep neural networks

Deep feedforward networks initialized along the edge of chaos exhibit exponentially superior training ability as quantified by maximum trainable depth. In this work, we explore the effect of saturation of the tanh activation function along the edge of chaos. In particular, we determine the line of uniformity in phase space along which the post-activation distribution has maximum entropy. This line intersects the edge of chaos, and indicates the regime beyond which saturation of the activation function begins to impede training efficiency. Our results suggest that initialization along the edge of chaos is a necessary but not sufficient condition for optimal trainability.

cs.LG

The edge of chaos: quantum field theory and deep neural networks

We explicitly construct the quantum field theory corresponding to a general class of deep neural networks encompassing both recurrent and feedforward architectures. We first consider the mean-field theory (MFT) obtained as the leading saddlepoint in the action, and derive the condition for criticality via the largest Lyapunov exponent. We then compute the loop corrections to the correlation function in a perturbative expansion in the ratio of depth $T$ to width $N$, and find a precise analogy with the well-studied $O(N)$ vector model, in which the variance of the weight initializations plays the role of the 't Hooft coupling. In particular, we compute both the $\mathcal{O}(1)$ corrections quantifying fluctuations from typicality in the ensemble of networks, and the subleading $\mathcal{O}(T/N)$ corrections due to finite-width effects. These provide corrections to the correlation length that controls the depth to which information can propagate through the network, and thereby sets the scale at which such networks are trainable by gradient descent. Our analysis provides a first-principles approach to the rapidly emerging NN-QFT correspondence, and opens several interesting avenues to the study of criticality in deep neural networks.

hep-th

Towards quantifying information flows: relative entropy in deep neural networks and the renormalization group

We investigate the analogy between the renormalization group (RG) and deep neural networks, wherein subsequent layers of neurons are analogous to successive steps along the RG. In particular, we quantify the flow of information by explicitly computing the relative entropy or Kullback-Leibler divergence in both the one- and two-dimensional Ising models under decimation RG, as well as in a feedforward neural network as a function of depth. We observe qualitatively identical behavior characterized by the monotonic increase to a parameter-dependent asymptotic value. On the quantum field theory side, the monotonic increase confirms the connection between the relative entropy and the c-theorem. For the neural networks, the asymptotic behavior may have implications for various information maximization methods in machine learning, as well as for disentangling compactness and generalizability. Furthermore, while both the two-dimensional Ising model and the random neural networks we consider exhibit non-trivial critical points, the relative entropy appears insensitive to the phase structure of either system. In this sense, more refined probes are required in order to fully elucidate the flow of information in these models.

hep-th

Hydrodynamic attractors in phase space

Hydrodynamic attractors have recently gained prominence in the context of early stages of ultra-relativistic heavy-ion collisions at the RHIC and LHC. We critically examine the existing ideas on this subject from a phase space point of view. In this picture the hydrodynamic attractor can be seen as a special case of the more general phenomenon of dynamical dimensionality reduction of phase space regions. We quantify this using Principal Component Analysis. Furthermore, we adapt the well known slow-roll approximation to this setting. These techniques generalize easily to higher dimensional phase spaces, which we illustrate by a preliminary analysis of a dataset describing the evolution of a 5-dimensional manifold of initial conditions immersed in a 16-dimensional representation of the phase space of the Boltzmann kinetic equation in the relaxation time approximation.

hep-th

On the efficacy of virtual seminars

During the SARS-CoV-2 pandemic, theoretical high-energy physics, and likely also the majority of other disciplines, are seeing a surge of virtual seminars as a primary means for scientific exchange. In this brief article, we highlight some compelling benefits of virtualizing research talks, and argue for why virtual seminars should continue even after the pandemic. Based on our extensive experience on running online talks, we also summarize some basic guidelines on organizing virtual seminars, and suggest some directions in which they could evolve.

physics.ed-ph

Information geometry in quantum field theory: lessons from simple examples

Motivated by the increasing connections between information theory and high-energy physics, particularly in the context of the AdS/CFT correspondence, we explore the information geometry associated to a variety of simple systems. By studying their Fisher metrics, we derive some general lessons that may have important implications for the application of information geometry in holography. We begin by demonstrating that the symmetries of the physical theory under study play a strong role in the resulting geometry, and that the appearance of an AdS metric is a relatively general feature. We then investigate what information the Fisher metric retains about the physics of the underlying theory by studying the geometry for both the classical 2d Ising model and the corresponding 1d free fermion theory, and find that the curvature diverges precisely at the phase transition on both sides. We discuss the differences that result from placing a metric on the space of theories vs. states, using the example of coherent free fermion states. We compare the latter to the metric on the space of coherent free boson states and show that in both cases the metric is determined by the symmetries of the corresponding density matrix. We also clarify some misconceptions in the literature pertaining to different notions of flatness associated to metric and non-metric connections, with implications for how one interprets the curvature of the geometry. Our results indicate that in general, caution is needed when connecting the AdS geometry arising from certain models with the AdS/CFT correspondence, and seek to provide a useful collection of guidelines for future progress in this exciting area.

hep-th

Path integral optimization as circuit complexity

Early efforts to understand complexity in field theory have primarily employed a geometric approach based on the concept of circuit complexity in quantum information theory. In a parallel vein, it has been proposed that certain deformations of the Euclidean path integral that prepares a given operator or state may provide an alternative definition, whose connection to the standard notion of complexity is less apparent. In this letter, we bridge the gap between these two proposals in two-dimensional conformal field theories, by explicitly showing how the latter approach from path integral optimization may be given a concrete realization within the standard gate counting framework. In particular, we show that when the background geometry is deformed by a Weyl rescaling, a judicious gate counting allows one to recover the Liouville action as a particular choice within a more general class of cost functions.

hep-th

Comments on black hole interiors and modular inclusions

We show how the traversable wormhole induced by a double-trace deformation of the thermofield double state can be understood as a modular inclusion of the algebras of exterior operators. The effect of this deformation is the creation of a new region of spacetime deep in the bulk, corresponding to a non-trivial center between the left and right algebras. This set-up provides a precise framework for investigating how black hole interiors are encoded in the CFT. In particular, we use modular theory to demonstrate that state dependence is an inevitable feature of any attempt to represent operators behind the horizon. Building on this geometrical structure, we propose that modular inclusions may provide a more precise means of investigating the nascent relationship between entanglement and geometry in the context of the emergent spacetime paradigm.

hep-th

Complexity as a novel probe of quantum quenches: universal scalings and purifications

We apply the recently developed notion of complexity for field theory to a quantum quench through a critical point in 1+1 dimensions. We begin with a toy model consisting of a quantum harmonic oscillator, and show that complexity exhibits universal scalings in both the slow and fast quench regimes. We then generalize our results to a 1-dimensional harmonic chain, and show that preservation of these scaling behaviours in free field theory depends on the choice of norm. Applying our set-up to the case of two oscillators, we quantify the complexity of purification associated to a subregion, and demonstrate that complexity is capable of probing features to which the entanglement entropy is insensitive. We find that the complexity of subregions is subadditive, and comment on potential implications for holography.

hep-th

Complexity and entanglement for thermofield double states

Motivated by holographic complexity proposals as novel probes of black hole spacetimes, we explore circuit complexity for thermofield double (TFD) states in free scalar quantum field theories using the Nielsen approach. For TFD states at t = 0, we show that the complexity of formation is proportional to the thermodynamic entropy, in qualitative agreement with holographic complexity proposals. For TFD states at t > 0, we demonstrate that the complexity evolves in time and saturates after a time of the order of the inverse temperature. The latter feature, which is in contrast with the results of holographic proposals, is due to the Gaussian nature of the TFD state of the free bosonic QFT. A novel technical aspect of our work is framing complexity calculations in the language of covariance matrices and the associated symplectic transformations, which provide a natural language for dealing with Gaussian states. Furthermore, for free QFTs in 1+1 dimension, we compare the dynamics of circuit complexity with the time dependence of the entanglement entropy for simple bipartitions of TFDs. We relate our results for the entanglement entropy to previous studies on non-equilibrium entanglement evolution following quenches. We also present a new analytic derivation of a logarithmic contribution due to the zero momentum mode in the limit of vanishing mass for a subsystem containing a single degree of freedom on each side of the TFD and argue why a similar logarithmic growth should be present for larger subsystems.

hep-th

Black holes and quantum entanglement

The black hole information paradox is a contradiction between fundamental principles which has puzzled physicists for over forty years. The crux of the problem lies in an assumption about the structure of entanglement across the event horizon, namely, that the Hilbert space factorizes. While valid in quantum mechanics, this fails drastically in quantum field theory, and hence a deeper understanding of entanglement is required if further progress is to be made. Recently, ideas from algebraic quantum field theory have provided new insight into this issue, and show promise for elucidating the connection between entropy and horizons that underlies black hole thermodynamics.

physics.pop-ph