SearcharxivSearch

arXiv subjects

Thomas R. Harvey

Publications and source records attributed to Thomas R. Harvey.

17 recordsLinked to original sources

Calculating Axion-Matter Couplings in String Theory

In this work we calculate, for the first time, the dominant perturbative contributions to axion--matter couplings explicitly in string theory. Concretely, we compute these couplings in Calabi--Yau compactifications of heterotic string theory where the vector bundle is a sum of line bundles. They are expressed as overlap integrals requiring the Ricci-flat metric, the Hermitian Yang-Mills bundle metric, and the matter field harmonic forms. We compute these by solving the corresponding coupled differential equations with neural networks. Our results establish a characteristic hierarchy of four-dimensional couplings: axions couple to fermions as $\mathcal{L} \supset c_ψ\,\frac{\partial_μa}{f_a}\, ψ^\dagger \barσ^μψ$ with $c_ψ$ an $\mathcal{O}(1)$ function of Kähler moduli for the model-dependent axions, while the corresponding coupling for the model-independent axion is suppressed by $α_{\rm YM}$ at the ultraviolet matching scale. The model-dependent axions thus couple to fermions similarly as DFSZ-like four-dimensional axions, whereas the model-independent axion's couplings are loop-suppressed as in KSVZ-like constructions. The methods developed here apply directly to the large class of heterotic line bundle models with realistic spectra, and can be adapted to type II constructions.

hep-th

What Neural Network Field Theory Can and Cannot Realise on a Computer

One aim of neural network field theory is to put a quantum or effective field theory on a computer, with the network ensemble itself as the theory. We ask how far that aim can be pushed for a function class regular enough to be computed with. Our main result is a no-go theorem with assumptions that hold for standard network architectures. We use it to separate four versions of neural network field theory, according to whether the defining object is the finite width ensemble or its infinite width limit, and whether the target we want to compute is a quantum or an effective field theory. Neither finite width interpretation is straightforwardly consistent. For finite width ensembles with finite variance at each point, the QFT interpretation fails reflection positivity, while the EFT interpretation establishes no scale separation by which the positivity violation can be placed outside its domain of validity. Of the two limit versions, one can be simulated in full and the other only in part, as only its smeared correlators are computable with a controlled error. As such, at the level of a controlled numerical computation, the QFT and EFT versions cannot be distinguished. One dimension escapes the obstruction, yet reflection positivity is shown to still fail there at every finite width for the cosine network. Two escapes from the theorem remain, giving up either finite variance at a point or exact rotation invariance, and we discuss both of these possibilities.

hep-th

Naturalness and Fisher Information

Fine-tuning and naturalness, the sensitivity of low-energy observables to small changes in the fundamental parameters of a theory, are cornerstones of physics beyond the Standard Model. We propose a new measure of fine-tuning based on information theory. To each point in parameter space we associate a probability distribution over observables. Divergence measures encode the sensitivity of observables to model parameters and determine a Riemannian metric on parameter space. By Chentsov's theorem, the physically motivated metric is the Fisher information metric, up to scaling. We propose a rescaled fine-tuning matrix $\mathcal{F}_{ij}$ derived from the Fisher information matrix, whose non-zero eigenvalues serve as our measure of fine-tuning. When the number of observables exceeds the number of parameters, $\mathcal{F}_{ij}$ admits a natural geometric interpretation as the pullback of the Euclidean metric from observable space to the submanifold of admissible predictions, with large eigenvalues corresponding to highly stretched directions and indicative of fine-tuning. Our measure reproduces the familiar Barbieri--Giudice criterion as a special case, while generalising it to multiple correlated parameters. We illustrate its behaviour on dimensional transmutation, the Wilson--Fisher fixed point, a simple model of the hierarchy problem, and the electron Yukawa coupling, finding agreement with physical intuition in each case.

hep-th

Sven: Singular Value Descent as a Computationally Efficient Natural Gradient Method

We introduce Sven (Singular Value dEsceNt), a new optimization algorithm for neural networks that exploits the natural decomposition of loss functions into a sum over individual data points, rather than reducing the full loss to a single scalar before computing a parameter update. Sven treats each data point's residual as a separate condition to be satisfied simultaneously, using the Moore-Penrose pseudoinverse of the loss Jacobian to find the minimum-norm parameter update that best satisfies all conditions at once. In practice, this pseudoinverse is approximated via a truncated singular value decomposition, retaining only the $k$ most significant directions and incurring a computational overhead of only a factor of $k$ relative to stochastic gradient descent. This is in comparison to traditional natural gradient methods, which scale as the square of the number of parameters. We show that Sven can be understood as a natural gradient method generalized to the over-parametrized regime, recovering natural gradient descent in the under-parametrized limit. On regression tasks, Sven significantly outperforms standard first-order methods including Adam, converging faster and to a lower final loss, while remaining competitive with LBFGS at a fraction of the wall-time cost. We discuss the primary challenge to scaling, namely memory overhead, and propose mitigation strategies. Beyond standard machine learning benchmarks, we anticipate that Sven will find natural application in scientific computing settings where custom loss functions decompose into several conditions.

cs.LG

The Optimiser Hidden in Plain Sight: Training with the Loss Landscape's Induced Metric

We present a class of novel optimisers for training neural networks that makes use of the Riemannian metric naturally induced when the loss landscape is embedded in higher-dimensional space. This is the same metric that underlies common visualisations of loss landscapes. By taking this geometric perspective literally and using the induced metric, we develop a new optimiser and compare it to existing methods, namely: SGD, Adam, AdamW, and Muon, across a range of tasks and architectures. Empirically, we conclude that this new class of optimisers is highly effective in low dimensional examples, and provides slight improvement over state-of-the-art methods for training neural networks. These new optimisers have theoretically desirable properties. In particular, the effective learning rate is automatically decreased in regions of high curvature acting as a smoothed out form of gradient clipping. Similarly, one variant of these optimisers can also be viewed as inducing an effective scheduled learning rate and decoupled weight decay is the natural choice from our geometric perspective. The basic method can be used to modify any existing preconditioning method. The new optimiser has a computational complexity comparable to that of Adam.

cs.LG

Fermion Masses and Mixing in String-Inspired Models

We study a class of supersymmetric Froggatt-Nielsen (FN) models with multiple U(1) symmetries and Standard Model (SM) singlets inspired by heterotic string compactifications on Calabi-Yau threefolds. The string-theoretic origin imposes a particular charge pattern on the SM fields and FN singlets, dividing the latter into perturbative and non-perturbative types. Employing systematic and heuristic search strategies, such as genetic algorithms, we identify charge assignments and singlet VEVs that replicate the observed mass and mixing hierarchies in the quark sector, and subsequently refine the Yukawa matrix coefficients to accurately match the observed values for the Higgs VEV, the quark and charged lepton masses and the CKM matrix. This bottom-up approach complements top-down string constructions and our results demonstrate that string FN models possess a sufficiently rich structure to account for flavour physics. On the other hand, the limited number of distinct viable charge patterns identified here indicates that flavour physics imposes tight constraints on string theory models, adding new constraints on particle spectra that are essential for achieving a realistic phenomenology.

hep-th

Symbolic Regression with Multimodal Large Language Models and Kolmogorov Arnold Networks

We present a novel approach to symbolic regression using vision-capable large language models (LLMs) and the ideas behind Google DeepMind's Funsearch. The LLM is given a plot of a univariate function and tasked with proposing an ansatz for that function. The free parameters of the ansatz are fitted using standard numerical optimisers, and a collection of such ansätze make up the population of a genetic algorithm. Unlike other symbolic regression techniques, our method does not require the specification of a set of functions to be used in regression, but with appropriate prompt engineering, we can arbitrarily condition the generative step. By using Kolmogorov Arnold Networks (KANs), we demonstrate that ``univariate is all you need'' for symbolic regression, and extend this method to multivariate functions by learning the univariate function on each edge of a trained KAN. The combined expression is then simplified by further processing with a language model.

cs.LG

Generative Modeling for Mathematical Discovery

We present a new implementation of the LLM-driven genetic algorithm {\it funsearch}, whose aim is to generate examples of interest to mathematicians and which has already had some success in problems in extremal combinatorics. Our implementation is designed to be useful in practice for working mathematicians; it does not require expertise in machine learning or access to high-performance computing resources. Applying {\it funsearch} to a new problem involves modifying a small segment of Python code and selecting a large language model (LLM) from one of many third-party providers. We benchmarked our implementation on three different problems, obtaining metrics that may inform applications of {\it funsearch} to new problems. Our results demonstrate that {\it funsearch} successfully learns in a variety of combinatorial and number-theoretic settings, and in some contexts learns principles that generalize beyond the problem originally trained on.

cs.LG

Not So Flat Metrics

In order to be in control of the $α'$ derivative expansion, geometric string compactifications are understood in the context of a large volume approximation. In this letter, we consider the reduction of these higher derivative terms, and propose an improved estimate on the large volume approximation using numerical Calabi-Yau metrics obtained via machine learning methods. Further to this, we consider the $α'^3$ corrections to numerical Calabi-Yau metrics in the context of IIB string theory. This correction represents one of several important contributions for realistic string compactifications -- alongside, for example, the backreaction of fluxes and local sources -- all of which have important consequences for string phenomenology. As a simple application of the corrected metric, we compute the change to the spectrum of the scalar Laplacian.

hep-th

Computation of Quark Masses from String Theory

We present a numerical computation, based on neural network techniques, of the physical Yukawa couplings in a heterotic string theory compactification on a smooth Calabi-Yau threefold with non-standard embedding. The model belongs to a large class of heterotic line bundle models that have previously been identified and whose low-energy spectrum precisely matches that of the MSSM plus fields uncharged under the Standard Model group. The relevant quantities for the calculation, that is, the Ricci-flat Calabi-Yau metric, the Hermitian Yang-Mills bundle metrics and the harmonic bundle-valued forms, are all computed by training suitable neural networks. For illustration, we consider a one-parameter family in complex structure moduli space. The computation at each point along this locus takes about half a day on a single twelve-core CPU. Our results for the Yukawa couplings are estimated to be within 10% of the expected analytic result. We find that the effect of the matter field normalisation can be significant and can contribute towards generating hierarchical couplings. We also demonstrate that a zeroth order, semi-analytic calculation, based on the Fubini-Study metric and its counterparts for the bundle metric and the bundle-valued forms, leads to roughly correct results, about 25% away from the numerical ones. The method can be applied to other heterotic line bundle models and generalised to other constructions, including to F-theory models.

hep-th

Spatially Homogeneous Universes with Late-Time Anisotropy

The cosmological principle asserts that on sufficiently large scales the Universe is homogeneous and isotropic on spatial slices. To deviate from this principle requires a departure from the FLRW ansatz. In this paper we analyze the cosmological evolution of two spatially homogeneous but anisotropic universes, namely the spatially closed Kantowski-Sachs Universe and the open axisymmetric Bianchi type III Universe. These models are characterized by two scale factors and we study their evolution in universes with radiation, matter and a cosmological constant. In all cases, the two scale factors evolve differently and this anisotropy leads to a lensing effect in the propagation of light. We derive explicit formulae for computing redshifts, angular diameter distances and luminosity distances and discuss the predictions of these models in relation to observations for type Ia supernovae and the CMB. We comment on the possibility of explaining the observed luminosity distance plot for type Ia supernovae within the context of cosmologies featuring late-time anisotropy and a vanishing cosmological constant.

astro-ph.CO

Enumerating Calabi-Yau Manifolds: Placing bounds on the number of diffeomorphism classes in the Kreuzer-Skarke list

The diffeomorphism class of simply-connected smooth Calabi-Yau threefolds with torsion-free cohomology is determined via certain basic topological invariants: the Hodge numbers, the triple intersection form, and the second Chern class. In the present paper, we shed some light on this classification by placing bounds on the number of diffeomorphism classes present in the set of smooth Calabi-Yau threefolds constructed from the Kreuzer-Skarke list of reflexive polytopes up to Picard number six. The main difficulty arises from the comparison of triple intersection numbers and divisor integrals of the second Chern class up to basis transformations. By using certain basis-independent invariants, some of which appear here for the first time, we are able to place lower bounds on the number of classes. Upper bounds are obtained by explicitly identifying basis transformations, using constraints related to the index of line bundles. Extrapolating our results, we conjecture that the favourable entries of the Kreuzer-Skarke list of reflexive polytopes leads to some $10^{400}$ diffeomorphically distinct Calabi-Yau threefolds.

hep-th

Decoding Nature with Nature's Tools: Heterotic Line Bundle Models of Particle Physics with Genetic Algorithms and Quantum Annealing

The string theory landscape may include a multitude of ultraviolet embeddings of the Standard Model, but identifying these has proven difficult due to the enormous number of available string compactifications. Genetic Algorithms (GAs) represent a powerful class of discrete optimisation techniques that can efficiently deal with the immensity of the string landscape, especially when enhanced with input from quantum annealers. In this letter we focus on geometric compactifications of the $E_8\times E_8$ heterotic string theory compactified on smooth Calabi-Yau threefolds with Abelian bundles. We make use of analytic formulae for bundle-valued cohomology to impose the entire range of spectrum requirements, something that has not been possible so far. For manifolds with a relatively low number of Kahler parameters we compare the GA search results with results from previous systematic scans, showing that GAs can find nearly all the viable solutions while visiting only a tiny fraction of the solution space. Moreover, we carry out GA searches on manifolds with a larger numbers of Kahler parameters where systematic searches are not feasible.

hep-th

Cosmic Inflation and Genetic Algorithms

Large classes of standard single-field slow-roll inflationary models consistent with the required number of e-folds, the current bounds on the spectral index of scalar perturbations, the tensor-to-scalar ratio, and the scale of inflation can be efficiently constructed using genetic algorithms. The setup is modular and can be easily adapted to include further phenomenological constraints. A semi-comprehensive search for sextic polynomial potentials results in roughly O(300,000) viable models for inflation. The analysis of this dataset reveals a preference for models with a tensor-to-scalar ratio in the range 0.0001 < r < 0.0004. We also consider potentials that involve cosine and exponential terms. In the last part we explore more complex methods of search relying on reinforcement learning and genetic programming. While reinforcement learning proves more difficult to use in this context, the genetic programming approach has the potential to uncover a multitude of viable inflationary models with new functional forms.

hep-th

String Model Building, Reinforcement Learning and Genetic Algorithms

We investigate reinforcement learning and genetic algorithms in the context of heterotic Calabi-Yau models with monad bundles. Both methods are found to be highly efficient in identifying phenomenologically attractive three-family models, in cases where systematic scans are not feasible. For monads on the bi-cubic Calabi-Yau either method facilitates a complete search of the environment and leads to similar sets of previously unknown three-family models.

hep-th

Evolving Heterotic Gauge Backgrounds: Genetic Algorithms versus Reinforcement Learning

The immensity of the string landscape and the difficulty of identifying solutions that match the observed features of particle physics have raised serious questions about the predictive power of string theory. Modern methods of optimisation and search can, however, significantly improve the prospects of constructing the standard model in string theory. In this paper we scrutinise a corner of the heterotic string landscape consisting of compactifications on Calabi-Yau three-folds with monad bundles and show that genetic algorithms can be successfully used to generate anomaly-free supersymmetric SO(10) GUTs with three families of fermions that have the right ingredients to accommodate the standard model. We compare this method with reinforcement learning and find that the two methods have similar efficacy but somewhat complementary characteristics.

hep-th

Heterotic String Model Building with Monad Bundles and Reinforcement Learning

We use reinforcement learning as a means of constructing string compactifications with prescribed properties. Specifically, we study heterotic SO(10) GUT models on Calabi-Yau three-folds with monad bundles, in search of phenomenologically promising examples. Due to the vast number of bundles and the sparseness of viable choices, methods based on systematic scanning are not suitable for this class of models. By focusing on two specific manifolds with Picard numbers two and three, we show that reinforcement learning can be used successfully to explore monad bundles. Training can be accomplished with minimal computing resources and leads to highly efficient policy networks. They produce phenomenologically promising states for nearly 100% of episodes and within a small number of steps. In this way, hundreds of new candidate standard models are found.

hep-th