SearcharxivSearch

arXiv subjects

Kit Fraser-Taliente

Publications and source records attributed to Kit Fraser-Taliente.

9 recordsLinked to original sources

The $T^{\mu\nu}$ of the conformal scalars

We construct the unique primary energy-momentum tensor $T^{\mu\nu}$ for the conformal free scalar with scaling dimension $\Delta=d/2-\zeta$ as a sum of Gegenbauer polynomials. For integer $\zeta$, the sum truncates at order $\zeta$, compactly reproducing all known results; for the nonlocal case of real $\zeta$, it is an infinite sum, with a two-parameter extension that reflects the nonuniqueness of the nonlocal geometric coupling. We find $T^{\mu\nu}$ by imposing off-shell conservation and tracelessness, and then directly solving the primary condition in momentum space. In the integer $\zeta$ case, we reproduce the known two-point function, and confirm the match with the $T^{\mu\nu}$ computed from Juhl's formulae for the GJMS operators (the Weyl-covariant upgrades of $(-\partial^2)^\zeta$), an equality following from the descent of Weyl covariance to conformal invariance.

hep-th

Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers

Large language model (LLM) activations are notoriously difficult to understand, with most existing techniques using complex, specialized methods for interpreting them. Recent work has proposed a simpler approach known as LatentQA: training LLMs to directly accept LLM activations as inputs and answer arbitrary questions about them in natural language. However, prior work has focused on narrow task settings for both training and evaluation. In this paper, we instead take a generalist perspective. We evaluate LatentQA-trained models, which we call Activation Oracles (AOs), in far out-of-distribution settings and examine how performance scales with training data diversity. We find that AOs can recover information fine-tuned into a model (e.g., biographical knowledge or malign propensities) that does not appear in the input text, despite never being trained with activations from a fine-tuned model. Our main evaluations are four downstream tasks where we can compare to prior white- and black-box techniques. We find that even narrowly-trained LatentQA models can generalize well, and that adding additional training datasets (such as classification tasks and a self-supervised context prediction task) yields consistent further improvements. Our best AOs match or exceed white-box baselines on all four tasks and the best overall baseline on 3 of 4. These results suggest that diversified training to answer natural-language queries imparts a general capability to verbalize information about LLM activations.

cs.CL

Inverse Scaling in Test-Time Compute

We construct evaluation tasks where extending the reasoning length of Large Reasoning Models (LRMs) deteriorates performance, exhibiting an inverse scaling relationship between test-time compute and accuracy. Our evaluation tasks span four categories: simple counting tasks with distractors, regression tasks with spurious features, deduction tasks with constraint tracking, and advanced AI risks. We identify five distinct failure modes when models reason for longer: 1) Claude models become increasingly distracted by irrelevant information; 2) OpenAI o-series models resist distractors but overfit to problem framings; 3) models shift from reasonable priors to spurious correlations; 4) all models show difficulties in maintaining focus on complex deductive tasks; and 5) extended reasoning may amplify concerning behaviors, with Claude Sonnet 4 showing increased expressions of self-preservation. These findings suggest that while test-time compute scaling remains promising for improving model capabilities, it may inadvertently reinforce problematic reasoning patterns. Our results demonstrate the importance of evaluating models across diverse reasoning lengths to identify and address these failure modes in LRMs.

cs.AI

Symbolic Regression with Multimodal Large Language Models and Kolmogorov Arnold Networks

We present a novel approach to symbolic regression using vision-capable large language models (LLMs) and the ideas behind Google DeepMind's Funsearch. The LLM is given a plot of a univariate function and tasked with proposing an ansatz for that function. The free parameters of the ansatz are fitted using standard numerical optimisers, and a collection of such ans\"atze make up the population of a genetic algorithm. Unlike other symbolic regression techniques, our method does not require the specification of a set of functions to be used in regression, but with appropriate prompt engineering, we can arbitrarily condition the generative step. By using Kolmogorov Arnold Networks (KANs), we demonstrate that ``univariate is all you need'' for symbolic regression, and extend this method to multivariate functions by learning the univariate function on each edge of a trained KAN. The combined expression is then simplified by further processing with a language model.

cs.LG

Diffusion Models for Cayley Graphs

We review the problem of finding paths in Cayley graphs of groups and group actions, using the Rubik's cube as an example, and we list several more examples of significant mathematical interest. We then show how to formulate these problems in the framework of diffusion models. The exploration of the graph is carried out by the forward process, while finding the target nodes is done by the inverse backward process. This systematizes the discussion and suggests many generalizations. To improve exploration, we propose a ``reversed score'' ansatz which substantially improves over previous comparable algorithms.

cs.LG

Not So Flat Metrics

In order to be in control of the $\alpha'$ derivative expansion, geometric string compactifications are understood in the context of a large volume approximation. In this letter, we consider the reduction of these higher derivative terms, and propose an improved estimate on the large volume approximation using numerical Calabi-Yau metrics obtained via machine learning methods. Further to this, we consider the $\alpha'^3$ corrections to numerical Calabi-Yau metrics in the context of IIB string theory. This correction represents one of several important contributions for realistic string compactifications -- alongside, for example, the backreaction of fluxes and local sources -- all of which have important consequences for string phenomenology. As a simple application of the corrected metric, we compute the change to the spectrum of the scalar Laplacian.

hep-th

Fermion Masses and Mixing in String-Inspired Models

We study a class of supersymmetric Froggatt-Nielsen (FN) models with multiple U(1) symmetries and Standard Model (SM) singlets inspired by heterotic string compactifications on Calabi-Yau threefolds. The string-theoretic origin imposes a particular charge pattern on the SM fields and FN singlets, dividing the latter into perturbative and non-perturbative types. Employing systematic and heuristic search strategies, such as genetic algorithms, we identify charge assignments and singlet VEVs that replicate the observed mass and mixing hierarchies in the quark sector, and subsequently refine the Yukawa matrix coefficients to accurately match the observed values for the Higgs VEV, the quark and charged lepton masses and the CKM matrix. This bottom-up approach complements top-down string constructions and our results demonstrate that string FN models possess a sufficiently rich structure to account for flavour physics. On the other hand, the limited number of distinct viable charge patterns identified here indicates that flavour physics imposes tight constraints on string theory models, adding new constraints on particle spectra that are essential for achieving a realistic phenomenology.

hep-th

Computation of Quark Masses from String Theory

We present a numerical computation, based on neural network techniques, of the physical Yukawa couplings in a heterotic string theory compactification on a smooth Calabi-Yau threefold with non-standard embedding. The model belongs to a large class of heterotic line bundle models that have previously been identified and whose low-energy spectrum precisely matches that of the MSSM plus fields uncharged under the Standard Model group. The relevant quantities for the calculation, that is, the Ricci-flat Calabi-Yau metric, the Hermitian Yang-Mills bundle metrics and the harmonic bundle-valued forms, are all computed by training suitable neural networks. For illustration, we consider a one-parameter family in complex structure moduli space. The computation at each point along this locus takes about half a day on a single twelve-core CPU. Our results for the Yukawa couplings are estimated to be within 10% of the expected analytic result. We find that the effect of the matter field normalisation can be significant and can contribute towards generating hierarchical couplings. We also demonstrate that a zeroth order, semi-analytic calculation, based on the Fubini-Study metric and its counterparts for the bundle metric and the bundle-valued forms, leads to roughly correct results, about 25% away from the numerical ones. The method can be applied to other heterotic line bundle models and generalised to other constructions, including to F-theory models.

hep-th

Enumerating Calabi-Yau Manifolds: Placing bounds on the number of diffeomorphism classes in the Kreuzer-Skarke list

The diffeomorphism class of simply-connected smooth Calabi-Yau threefolds with torsion-free cohomology is determined via certain basic topological invariants: the Hodge numbers, the triple intersection form, and the second Chern class. In the present paper, we shed some light on this classification by placing bounds on the number of diffeomorphism classes present in the set of smooth Calabi-Yau threefolds constructed from the Kreuzer-Skarke list of reflexive polytopes up to Picard number six. The main difficulty arises from the comparison of triple intersection numbers and divisor integrals of the second Chern class up to basis transformations. By using certain basis-independent invariants, some of which appear here for the first time, we are able to place lower bounds on the number of classes. Upper bounds are obtained by explicitly identifying basis transformations, using constraints related to the index of line bundles. Extrapolating our results, we conjecture that the favourable entries of the Kreuzer-Skarke list of reflexive polytopes leads to some $10^{400}$ diffeomorphically distinct Calabi-Yau threefolds.

hep-th