SearcharxivSearch

arXiv subjects

James Halverson

Publications and source records attributed to James Halverson.

At least 19 recordsLinked to original sources

Spinning Conformal Correlators from Neural Networks

We construct spinning conformal fields from neural networks and the embedding formalism, computing their two-, three- and four-point functions in examples, building on scalar conformal field techniques introduced in \cite{Halverson:2024axc}. For a particular ensemble of i.i.d. neurons we recover the 4d Maxwell CFT in the infinite-width limit.

hep-th

Pre-Strings Lectures on Artificial Intelligence

These notes are based on six lectures given over three days at the Pre-Strings 2026 school in Shanghai. Day 1 develops neural network essentials, organized around the expressivity, statistics, and dynamics of neural networks, presented with a field-theoretic lens. Day 2 develops a neural network approach to field theory (NN-FT), in which a field theory is defined by a network architecture and a density on its parameters, and surveys recent results. Examples include a universality theorem, a neural network realization of Liouville theory, famous string amplitudes, topological sectors and the Kosterlitz-Thouless transition, Ward identities and anomalies, and a new derivation of the critical dimension of the bosonic string. Day 3 turns the lens around and covers applied AI for string theory: agentic workflows that are changing how the other techniques are implemented, physics-informed neural networks and Calabi-Yau metrics, reinforcement learning and search in the string landscape and in knot theory, and interpretable supervised learning with an eye towards conjecture generation.

hep-th

A Supernova Constraint on F-theory

We study constraints on the quantum chromodynamics (QCD) axion in F-theory, a strongly coupled limit of string theory. We build models of QCD from compactifications to four dimensions on elliptic fibrations over toric threefolds $B_3$, characterized by an integer $N=h^{1,1}(B_3)$. The QCD axion mass increases with $N$, and we find that models with sufficiently high $N$ are inconsistent with observations of the neutrino burst from supernova 1987A, disfavoring large regions of the moduli space. Specifically, at least $95\%$ of models with $N \ge 8,791$ -- a regime that arguably contains the vast majority of known F-theory topologies -- have a QCD axion mass $m_{QCD}>15~meV$ and are thus constrained. This limit is independent of cosmology. We only consider weakly-curved threefolds, where $\alpha'$ corrections are plausibly negligible.

hep-th

Anomalies in Neural Network Field Theory

Neural network field theory (NN-FT) formulates field theory in terms of a network architecture and a density on its parameters. We derive Schwinger--Dyson equations and Ward identities in NN-FT and utilize them to study anomalies. The equations depend on a conserved parameter space current that characterizes symmetries and how they break. It is relevant even in non-local NN-FTs, but can recover local currents in the case of a local Lagrangian by an appropriate fiber-wise average. In machine learning, this formalism is applied to feedforward networks and the attention mechanism. In physics, we use this machinery to study $U(1)$ symmetry for a complex scalar, the scale anomaly in $4d$ massless $\phi^4$ theory, the Weyl anomaly for the bosonic string (including a new computation of the critical dimension), and examples involving discrete topological data, such as winding numbers and T-duality. Since the results are obtained in network parameter space rather than the standard field space, they represent a new way to understand symmetries in quantum field theories.

hep-th

Small Vacuum Energy and Tunneling in a Modified Bousso-Polchinski Model

We propose a simplified model for the cosmological constant in string theory flux vacua motivated by type IIB and F-theory compactifications. Relative to the Bousso-Polchinski model, small vacuum energy spacing occurs in thin wafers rather than thin shells. The model is applied to the entire Sch\"oller-Skarke database of Calabi-Yau fourfolds, which exhibit $532,600,483$ distinct sets of Hodge numbers. The overwhelming majority of those ($99.95\%$ percent for some choices of parameters) exhibit a vacuum energy spacing of~$10^{-120}$ in Planck units or smaller. Brown-Teitelboim membrane nucleation transitions can populate this landscape of flux vacua. In the thin-wall approximation, and ignoring gravitational corrections, we find that the bubble transitions are always dominated by giant leaps in flux space. The age of the universe places a bound on Calabi-Yau topology that is satisfied for the entire Sch\"oller-Skarke database.

hep-th

Topological Effects in Neural Network Field Theory

Neural network field theory formulates field theory as a statistical ensemble of fields defined by a network architecture and a density on its parameters. We extend the construction to topological settings via the inclusion of discrete parameters that label the topological quantum number. We recover the Berezinskii--Kosterlitz--Thouless transition, including the spin-wave critical line and the proliferation of vortices at high temperatures. We also verify the T-duality of the bosonic string, showing invariance under the exchange of momentum and winding on $S^1$, the transformation of the sigma model couplings according to the Buscher rules on constant toroidal backgrounds, the enhancement of the current algebra at self-dual radius, and non-geometric T-fold transition functions.

hep-th

Naturalness and Fisher Information

Fine-tuning and naturalness, the sensitivity of low-energy observables to small changes in the fundamental parameters of a theory, are cornerstones of physics beyond the Standard Model. We propose a new measure of fine-tuning based on information theory. To each point in parameter space we associate a probability distribution over observables. Divergence measures encode the sensitivity of observables to model parameters and determine a Riemannian metric on parameter space. By Chentsov's theorem, the physically motivated metric is the Fisher information metric, up to scaling. We propose a rescaled fine-tuning matrix $\mathcal{F}_{ij}$ derived from the Fisher information matrix, whose non-zero eigenvalues serve as our measure of fine-tuning. When the number of observables exceeds the number of parameters, $\mathcal{F}_{ij}$ admits a natural geometric interpretation as the pullback of the Euclidean metric from observable space to the submanifold of admissible predictions, with large eigenvalues corresponding to highly stretched directions and indicative of fine-tuning. Our measure reproduces the familiar Barbieri--Giudice criterion as a special case, while generalising it to multiple correlated parameters. We illustrate its behaviour on dimensional transmutation, the Wilson--Fisher fixed point, a simple model of the hierarchy problem, and the electron Yukawa coupling, finding agreement with physical intuition in each case.

hep-th

Universality of Neural Network Field Theory

We prove that any quantum field theory, or more generally any probability distribution over tempered distributions in $\mathbb{R}^d$, admits a neural network description with a countable infinity of parameters. As an example, we realize the $2d$ Liouville theory as a neural network and numerically compute the three-point function of vertex operators, finding agreement with the DOZZ formula.

hep-th

String Theory from Infinite Width Neural Networks

We realize bosonic string theory with ensembles of infinite width neural networks. The string tension is tuned by the variance of the output weights. The construction provides a new computation of the foundational Virasoro-Shapiro and Veneziano amplitudes as neural network correlators.

hep-th

F-theory Axiverse

We compute the couplings of Ramond-Ramond four-form axions in three ensembles of F-theory compactifications, with up to 181,200 axions. We work in the stretched K\"ahler cone, where $\alpha'$ corrections are plausibly controlled, and we use couplings to certain non-Abelian sectors as a proxy for couplings to photons. The axion masses, decay constants, and couplings to gauge sectors show striking universality across the ensembles. In particular, the axion-photon couplings grow with $h^{1,1}$, and models in our ensemble with $h^{1,1} \gtrsim$ 10,000 axions are in tension with helioscope constraints. Moreover, under mild assumptions about charged matter beyond the Standard Model, theories with $h^{1,1} \gtrsim$ 5,000 are in tension with Chandra measurements of X-ray spectra. This work is a first step toward understanding the phenomenology of quantum gravity theories with thousands of axions.

hep-th

Fermions and Supersymmetry in Neural Network Field Theories

We introduce fermionic neural network field theories via Grassmann-valued neural networks. Free theories are obtained by a generalization of the Central Limit Theorem to Grassmann variables. This enables the realization of the free Dirac spinor at infinite width and a four fermion interaction at finite width. Yukawa couplings are introduced by breaking the statistical independence of the output weights for the fermionic and bosonic fields. A large class of interacting supersymmetric quantum mechanics and field theory models are introduced by super-affine transformations on the input that realize a superspace formalism.

hep-th

Symbolic Regression with Multimodal Large Language Models and Kolmogorov Arnold Networks

We present a novel approach to symbolic regression using vision-capable large language models (LLMs) and the ideas behind Google DeepMind's Funsearch. The LLM is given a plot of a univariate function and tasked with proposing an ansatz for that function. The free parameters of the ansatz are fitted using standard numerical optimisers, and a collection of such ans\"atze make up the population of a genetic algorithm. Unlike other symbolic regression techniques, our method does not require the specification of a set of functions to be used in regression, but with appropriate prompt engineering, we can arbitrarily condition the generative step. By using Kolmogorov Arnold Networks (KANs), we demonstrate that ``univariate is all you need'' for symbolic regression, and extend this method to multivariate functions by learning the univariate function on each edge of a trained KAN. The combined expression is then simplified by further processing with a language model.

cs.LG

Learning Topological Invariance

Two geometric spaces are in the same topological class if they are related by certain geometric deformations. We propose machine learning methods that automate learning of topological invariance and apply it in the context of knot theory, where two knots are equivalent if they are related by ambient space isotopy. Specifically, given only the knot and no information about its topological invariants, we employ contrastive and generative machine learning techniques to map different representatives of the same knot class to the same point in an embedding vector space. An auto-regressive decoder Transformer network can then generate new representatives from the same knot class. We also describe a student-teacher setup that we use to interpret which known knot invariants are learned by the neural networks to compute the embeddings, and observe a strong correlation with the Goeritz matrix in all setups that we tested. We also develop an approach to resolving the Jones Unknot Conjecture by exploring the vicinity of the embedding space of the Jones polynomial near the locus where the unknots cluster, which we use to generate braid words with simple Jones polynomials.

math.GT

Quantum Mechanics and Neural Networks

We demonstrate that any Euclidean-time quantum mechanical theory may be represented as a neural network, ensured by the Kosambi-Karhunen-Lo\`eve theorem, mean-square path continuity, and finite two-point functions. The additional constraint of reflection positivity, which is related to unitarity, may be achieved by a number of mechanisms, such as imposing neural network parameter space splitting or the Markov property. Non-differentiability of the networks is related to the appearance of non-trivial commutators. Neural networks acting on Markov processes are no longer Markov, but still reflection positive, which facilitates the definition of deep neural network quantum systems. We illustrate these principles in several examples using numerical implementations, recovering classic quantum mechanical results such as Heisenberg uncertainty, non-trivial commutators, and the spectrum.

hep-th

Conversations and Deliberations: Non-Standard Cosmological Epochs and Expansion Histories

This document summarizes the discussions which took place during the PITT-PACC Workshop entitled "Non-Standard Cosmological Epochs and Expansion Histories," held in Pittsburgh, Pennsylvania, Sept. 5-7, 2024. Much like the non-standard cosmological epochs that were the subject of these discussions, the format of this workshop was also non-standard. Rather than consisting of a series of talks from participants, with each person presenting their own work, this workshop was instead organized around free-form discussion blocks, with each centered on a different overall theme and guided by a different set of Discussion Leaders. This document is not intended to serve as a comprehensive review of these topics, but rather as an informal record of the discussions that took place during the workshop, in the hope that the content and free-flowing spirit of these discussions may inspire new ideas and research directions.

astro-ph.CO

Conformal Fields from Neural Networks

We use the embedding formalism to construct conformal fields in $D$ dimensions, by restricting Lorentz-invariant ensembles of homogeneous neural networks in $(D+2)$ dimensions to the projective null cone. Conformal correlators may be computed using the parameter space description of the neural network. Exact four-point correlators are computed in a number of examples, and we perform a 4D conformal block decomposition that elucidates the spectrum. In some examples the analysis is facilitated by recent approaches to Feynman integrals. Generalized free CFTs are constructed using the infinite-width Gaussian process limit of the neural network, enabling a realization of the free boson. The extension to deep networks constructs conformal fields at each subsequent layer, with recursion relations relating their conformal dimensions and four-point functions. Numerical approaches are discussed.

hep-th

On the Generality and Persistence of Cosmological Stasis

Hierarchical decays of $N$ matter species to radiation may balance against Hubble expansion to yield stasis, a new phase of cosmological evolution with constant matter and radiation abundances. We analyze stasis with various machine learning techniques on the full $2N$-dimensional space of decay rates and abundances, which serve as inputs to the system of Boltzmann equations that governs the dynamics. We construct a differentiable Boltzmann solver to maximize the number of stasis $e$-folds $\mathcal{N}$. High-stasis configurations obtained by gradient ascent motivate log-uniform distributions on rates and abundances to accompany power-law distributions of previous works. We demonstrate that random configurations drawn from these families of distributions regularly exhibit many $e$-folds of stasis. We additionally use them as priors in a Bayesian analysis conditioned on stasis, using stochastic variational inference with normalizing flows to model the posterior. All three numerical analyses demonstrate the generality of stasis and point to a new model in which the rates and abundances are exponential in the species index. We show that the exponential model solves the exact stasis equations, is an attractor, and satisfies $\mathcal{N}\propto N$, exhibiting inflation-level $e$-folding with a relatively low number of species. This is contrasted with the $\mathcal{N}\propto \log(N)$ scaling of power-law models. Finally, we discuss implications for the emergent string conjecture and string axiverse.

astro-ph.CO

KAN: Kolmogorov-Arnold Networks

Inspired by the Kolmogorov-Arnold representation theorem, we propose Kolmogorov-Arnold Networks (KANs) as promising alternatives to Multi-Layer Perceptrons (MLPs). While MLPs have fixed activation functions on nodes ("neurons"), KANs have learnable activation functions on edges ("weights"). KANs have no linear weights at all -- every weight parameter is replaced by a univariate function parametrized as a spline. We show that this seemingly simple change makes KANs outperform MLPs in terms of accuracy and interpretability. For accuracy, much smaller KANs can achieve comparable or better accuracy than much larger MLPs in data fitting and PDE solving. Theoretically and empirically, KANs possess faster neural scaling laws than MLPs. For interpretability, KANs can be intuitively visualized and can easily interact with human users. Through two examples in mathematics and physics, KANs are shown to be useful collaborators helping scientists (re)discover mathematical and physical laws. In summary, KANs are promising alternatives for MLPs, opening opportunities for further improving today's deep learning models which rely heavily on MLPs.

cs.LG