SearcharxivSearch

arXiv subjects

David S. Berman

Publications and source records attributed to David S. Berman.

At least 19 recordsLinked to original sources

AI and the Research-Education Environment of Physics

In the current era of AI transforming the research-education environment of physics, variety of issues and concerns arise. The KITP program "Generative AI for High and Low Energy Physics'' offered a discussion session on this, and here presented is a summary of the opinions provided in the discussion. The material is formulated such that it can serve as a starting point for further discussions in readers' research community/institution/group.

physics.ed-ph

A path to natural language through tokenisation and transformers

Natural languages exhibit striking regularities in their statistical structure, including notably the emergence of Zipf's and Heaps' laws. Despite this, it remains broadly unclear how these properties relate to the modern tokenisation schemes used in contemporary transformer models. In this note, we analyse the information content (as measured by the Shannon entropy) of various corpora under the assumption of a Zipfian frequency distribution, and derive a closed-form expression for the slot entropy expectation value. We then empirically investigate how byte--pair encoding (BPE) transforms corpus statistics, showing that recursive applications of BPE drive token frequencies toward a Zipfian power law while inducing a characteristic growth pattern in empirical entropy. Utilizing the ability of transformers to learn context dependent token probability distributions, we train language models on corpora tokenised at varying BPE depths, revealing that the model predictive entropies increasingly agree with Zipf-derived predictions as the BPE depth increases. Attention-based diagnostics further indicate that deeper tokenisation reduces local token dependencies, bringing the empirical distribution closer to the weakly dependent (near IID) regime. Together, these results clarify how BPE acts not only as a compression mechanism but also as a statistical transform that reconstructs key informational properties of natural language.

cs.CL

Modeling financial time series with $\phi^{4}$ quantum field theory

We use a $\phi^{4}$ quantum field theory with inhomogeneous couplings and explicit symmetry-breaking to model an ensemble of financial time series from the S$\&$P 500 index. The continuum nature of the $\phi^4$ theory avoids the inaccuracies that occur in Ising-based models which require a discretization of the time series. We demonstrate this using the example of the 2008 global financial crisis. The $\phi^{4}$ quantum field theory is expressive enough to reproduce the higher-order statistics such as the market kurtosis, which can serve as an indicator of possible market shocks. Accurate reproduction of high kurtosis is absent in binarized models. Therefore Ising models, despite being widely employed in econophysics, are incapable of fully representing empirical financial data, a limitation not present in the generalization of the $\phi^{4}$ scalar field theory. We then investigate the scaling properties of the $\phi^{4}$ machine learning algorithm and extract exponents which govern the behavior of the learned couplings (or weights and biases in ML language) in relation to the number of stocks in the model. Finally, we use our model to forecast the price changes of the AAPL, MSFT, and NVDA stocks. We conclude by discussing how the $\phi^{4}$ scalar field theory could be used to build investment strategies and the possible intuitions that the QFT operations of dimensional compactification and renormalization can provide for financial modelling.

q-fin.ST

Grokking vs. Learning: Same Features, Different Encodings

Grokking typically achieves similar loss to ordinary, "steady", learning. We ask whether these different learning paths - grokking versus ordinary training - lead to fundamental differences in the learned models. To do so we compare the features, compressibility, and learning dynamics of models trained via each path in two tasks. We find that grokked and steadily trained models learn the same features, but there can be large differences in the efficiency with which these features are encoded. In particular, we find a novel "compressive regime" of steady training in which there emerges a linear trade-off between model loss and compressibility, and which is absent in grokking. In this regime, we can achieve compression factors 25x times the base model, and 5x times the compression achieved in grokking. We then track how model features and compressibility develop through training. We show that model development in grokking is task-dependent, and that peak compressibility is achieved immediately after the grokking plateau. Finally, novel information-geometric measures are introduced which demonstrate that models undergoing grokking follow a straight path in information space.

cs.LG

Curvature of an exotic 7-sphere

We study the geometry of the Gromoll-Meyer sphere, one of Milnor's exotic $7$-spheres. We focus on a Kaluza-Klein Ansatz, with a round $S^4$ as base space, unit $S^3$ as fibre, and $k=1,2$ $SU(2)$ instantons as gauge fields, where all quantities admit an elegant description in quaternionic language. The metric's moduli space coincides with the $k=1,2$ instantons' moduli space quotiented by the isometry of the base, plus an additional $\mathbb{R}^+$ factor corresponding to the radius of the base, $r$. We identify a "center" of the $k=2$ instanton moduli space with enhanced symmetry. This $k=2$ solution is used together with the maximally symmetric $k=1$ solution to obtain a metric of maximal isometry, $SO(3)\times O(2)$, and to explicitly compute its Ricci tensor. This allows us to put a bound on $r$ to ensure positive Ricci curvature, which implies various energy conditions for an $8$-dimensional static space-time. This construction then enables a concrete examination of the properties of the sectional curvature.

hep-th

NCoder -- A Quantum Field Theory approach to encoding data

In this paper we present a novel approach to interpretable AI inspired by Quantum Field Theory (QFT) which we call the NCoder. The NCoder is a modified autoencoder neural network whose latent layer is prescribed to be a subset of $n$-point correlation functions. Regarding images as draws from a lattice field theory, this architecture mimics the task of perturbatively constructing the effective action of the theory order by order in an expansion using Feynman diagrams. Alternatively, the NCoder may be regarded as simulating the procedure of statistical inference whereby high dimensional data is first summarized in terms of several lower dimensional summary statistics (here the $n$-point correlation functions), and subsequent out-of-sample data is generated by inferring the data generating distribution from these statistics. In this way the NCoder suggests a fascinating correspondence between perturbative renormalizability and the sufficiency of models. We demonstrate the efficacy of the NCoder by applying it to the generation of MNIST images, and find that generated images can be correctly classified using only information from the first three $n$-point functions of the image distribution.

hep-th

Bayesian Renormalization

In this note we present a fully information theoretic approach to renormalization inspired by Bayesian statistical inference, which we refer to as Bayesian Renormalization. The main insight of Bayesian Renormalization is that the Fisher metric defines a correlation length that plays the role of an emergent RG scale quantifying the distinguishability between nearby points in the space of probability distributions. This RG scale can be interpreted as a proxy for the maximum number of unique observations that can be made about a given system during a statistical inference experiment. The role of the Bayesian Renormalization scheme is subsequently to prepare an effective model for a given system up to a precision which is bounded by the aforementioned scale. In applications of Bayesian Renormalization to physical systems, the emergent information theoretic scale is naturally identified with the maximum energy that can be probed by current experimental apparatus, and thus Bayesian Renormalization coincides with ordinary renormalization. However, Bayesian Renormalization is sufficiently general to apply even in circumstances in which an immediate physical scale is absent, and thus provides an ideal approach to renormalization in data science contexts. To this end, we provide insight into how the Bayesian Renormalization scheme relates to existing methods for data compression and data generation such as the information bottleneck and the diffusion learning paradigm. We conclude by designing an explicit form of Bayesian Renormalization inspired by Wilson's momentum shell renormalization scheme in Quantum Field Theory. We apply this Bayesian Renormalization scheme to a simple Neural Network and verify the sense in which it organizes the parameters of the model according to a hierarchy of information theoretic importance.

hep-th

The Inverse of Exact Renormalization Group Flows as Statistical Inference

We build on the view of the Exact Renormalization Group (ERG) as an instantiation of Optimal Transport described by a functional convection-diffusion equation. We provide a new information theoretic perspective for understanding the ERG through the intermediary of Bayesian Statistical Inference. This connection is facilitated by the Dynamical Bayesian Inference scheme, which encodes Bayesian inference in the form of a one parameter family of probability distributions solving an integro-differential equation derived from Bayes' law. In this note, we demonstrate how the Dynamical Bayesian Inference equation is, itself, equivalent to a diffusion equation which we dub Bayesian Diffusion. Identifying the features that define Bayesian Diffusion, and mapping them onto the features that define the ERG, we obtain a dictionary outlining how renormalization can be understood as the inverse of statistical inference.

hep-th

Twisted Self-duality

We examine a generalisation of the usual self-duality equations for Yang-Mills theory when the colour space admits a non-trivial involution. This involution allows us to construct a non-trivial twist which may be combined with the Hodge star to form a twisted self-dual curvature. We will construct a simple example of twisted self-duality for $su(2) \oplus su(2)$ gauge theory along with its explicit solutions and then dimensionally reduce from four dimensions to obtain families of non-trivial non-linear equations in lower dimensions. This twisted self-duality constraint will be shown to arise in E_7 exceptional field theory through a Scherk-Schwarz reduction and we will show how an Eguchi-Hanson gravitational instanton also obeys the twisted self-duality condition.

hep-th

On the Dynamics of Inference and Learning

Statistical Inference is the process of determining a probability distribution over the space of parameters of a model given a data set. As more data becomes available this probability distribution becomes updated via the application of Bayes' theorem. We present a treatment of this Bayesian updating process as a continuous dynamical system. Statistical inference is then governed by a first order differential equation describing a trajectory or flow in the information geometry determined by a parametric family of models. We solve this equation for some simple models and show that when the Cramér-Rao bound is saturated the learning rate is governed by a simple $1/T$ power-law, with $T$ a time-like variable denoting the quantity of data. The presence of hidden variables can be incorporated in this setting, leading to an additional driving term in the resulting flow equation. We illustrate this with both analytic and numerical examples based on Gaussians and Gaussian Random Processes and inference of the coupling constant in the 1D Ising model. Finally we compare the qualitative behaviour exhibited by Bayesian flows to the training of various neural networks on benchmarked data sets such as MNIST and CIFAR10 and show how that for networks exhibiting small final losses the simple power-law is also satisfied.

cond-mat.dis-nn

Double copying Exceptional Field theories

We examine exceptional field theory through the lens of the generalised double copy formalism. This allows us to construct classical solutions in M-theory using a generalised Kerr-Schild ansatz and along the way indicates hints towards a single copy of M-theory. Based on a talk at the Nankai Symposium given by DSB.

hep-th

Machine Learning Calabi-Yau Hypersurfaces

We revisit the classic database of weighted-P4s which admit Calabi-Yau 3-fold hypersurfaces equipped with a diverse set of tools from the machine-learning toolbox. Unsupervised techniques identify an unanticipated almost linear dependence of the topological data on the weights. This then allows us to identify a previously unnoticed clustering in the Calabi-Yau data. Supervised techniques are successful in predicting the topological parameters of the hypersurface from its weights with an accuracy of R^2 > 95%. Supervised learning also allows us to identify weighted-P4s which admit Calabi-Yau hypersurfaces to 100% accuracy by making use of partitioning supported by the clustering behaviour.

hep-th

The Classical Double Copy for M-theory from a Kerr-Schild Ansatz for Exceptional Field Theory

We construct the classical double copy formalism for M-theory. This extends the current state of the art by including the three form potential of eleven dimensional supergravity along with the metric. The key for this extension is to construct a Kerr-Schild type Ansatz for exceptional field theory. This Kerr-Schild Ansatz then allows us to find the solutions of charged objects such as the membrane from a set of single copy fields. The exceptional field theory formalism then automatically produces the IIB Kerr-Schild ansatz allowing the construction of the single copy for the fields of IIB supergravity (with manifest $SL(2)$ symmetry).

hep-th

The single copy of the gravitational holonomy

The double copy is a well-established relationship between gravity and gauge theories. It relates perturbative scattering amplitudes as well as classical solutions, and recently there has been mounting evidence that it also applies to non-perturbative information. In this paper, we consider the holonomy properties of manifolds in gravity and prescribe a single copy of gravitational holonomy that differs from the holonomy in gauge theory. We discuss specific cases and give examples where the single copy holonomy group is reduced. Our results may prove useful in extending the classical double copy. We also clarify previous misconceptions in the literature regarding gravitational Wilson lines and holonomy.

hep-th

Double Field Theory and Geometric Quantisation

We examine various properties of double field theory and the doubled string sigma model in the context of geometric quantisation. In particular we look at T-duality as the symplectic transformation related to an alternative choice of polarisation in the construction of the quantum bundle for the string. Following this perspective we adopt a variety of techniques from geometric quantisation to study the doubled space. One application is the construction of the double coherent state that provides the shortest distance in any duality frame and a stringy deformed Fourier transform.

hep-th

Weyl doubling

We study a host of spacetimes where the Weyl curvature may be expressed algebraically in terms of an Abelian field strength. These include Type D spacetimes in four and higher dimensions which obey a simple quadratic relation between the field strength and the Weyl tensor, following the Weyl spinor double copy relation. However, we diverge from the usual double copy paradigm by taking the gauge fields to be in the curved spacetime as opposed to an auxiliary flat space. We show how for Gibbons-Hawking spacetimes with more than two centres a generalisation of the Weyl doubling formula is needed by including a derivative-dependent expression which is linear in the Abelian field strength. We also find a type of twisted doubling formula in a case of a manifold with Spin(7) holonomy in eight dimensions. For Einstein Maxwell theories where there is an independent gauge field defined on spacetime, we investigate how the gauge fields determine the Weyl spacetime curvature via a doubling formula. We first show that this occurs for the Reissner-Nordstrom metric in any dimension, and that this generalises to the electrically-charged Born-Infeld solutions. Finally, we consider brane systems in supergravity, showing that a similar doubling formula applies. This Weyl formula is based on the field strength of the p-form potential that minimally couples to the brane and the brane world volume Killing vectors.

hep-th

The Geometry, Branes and Applications of Exceptional Field Theory

This is a review of exceptional field theory: a generalisation of Kaluza-Klein theory that unifies the metric and $p$-form gauge field degrees of freedom of supergravity into a generalised or extended geometry, whose additional coordinates may be viewed as conjugate to brane winding modes. This unifies the maximal supergravities, treating their previously-hidden exceptional Lie symmetries as a fundamental geometric symmetry. Duality orbits of solutions simplify into single objects, that in many cases have simple geometric interpretations, for instance as wave or monopole-type solutions. It also provides a route to explore exotic or non-geometric aspects of M-theory, such as exotic branes, U-folds, and more novel sorts of non-Riemannian spaces.

hep-th

S-duality and the Double Copy

The double copy formalism provides an intriguing connection between gauge theories and gravity. It was first demonstrated in the perturbative context of scattering amplitudes but recently the formalism has been applied to exact classical solutions in gauge theories such as the monopole and instanton. In this paper we will investigate how duality symmetries in the gauge theory double copy to gravity and relate these to solution generating transformations and the action of $Sl(2,R)$ in general relativity.

hep-th