SearcharxivSearch

arXiv subjects

Fabian Ruehle

Publications and source records attributed to Fabian Ruehle.

At least 19 recordsLinked to original sources

Learning Topological Features of $\widehat Z$-invariants

Machine learning and data analysis techniques have recently emerged as powerful tools for identifying patterns and formulating conjectures in mathematical research, most notably in the field of low-dimensional topology. In this paper, we initiate a systematic approach to handling mathematical data structured as (truncated) infinite $q$-series, or equivalently, infinite series of integers. To apply this data analysis pipeline, we construct a comprehensive dataset of $\widehat{Z}$-invariants (homological blocks) for plumbed 3-manifolds. We demonstrate that neural networks can reliably extract essential topological information, such as homology class and underlying graph structure, directly from the $q$-series coefficients. A central feature of our methodology is a focus on interpretability; by contrasting local gradient sensitivity with global feature relevance, we reveal that the networks learn to bypass complex topological rules in favor of specific spectral and geometric proxies. Finally, we apply this pipeline to probe homology cobordism, discovering a high-accuracy predictive relationship between the $\widehat{Z}$-invariant exponents and the Heegaard Floer $d$-invariant (correction term). These results suggest that $\widehat{Z}$-invariants capture subtle geometric information regarding cobordism equivalences, warranting a new direction for the study of quantum invariants.

hep-th

Warped Numerical Calabi-Yau Metrics

We compute numerical warped Type IIB flux backgrounds on Calabi-Yau threefolds following the construction of Giddings, Kachru, and Polchinski. Using physics-informed neural networks, we approximate all three ingredients required by the GKP setup: the Ricci-flat Calabi-Yau metric, the harmonic (2,1)-forms representing the imaginary self-dual three-form flux, and the warp factor, which solves a sourced Poisson equation on the internal manifold. We apply our pipeline to the Dwork family of quintics for two different flux vacua, one near a conifold point and one away from it, the latter serving as a numerical cross-check. With these tools, we study the singular bulk problem, and find that for our benchmark point close to the conifold, approximately 0.5 percent of the total Calabi-Yau volume sits in the throat, and the warp factor is an order of magnitude larger as compared to the bulk. We also introduce several improvements to techniques used for numerical studies of CY metrics and quantities derived from them that might be of interest independently of our application. These include an improved point sampling algorithm that produces samples that are more uniform under the Calabi-Yau measure, a feature-engineered spectral network for the metric, multi-step physics-informed training, and a weighted Huber loss tailored to stiff PDEs with highly non-uniform sources.

hep-th

Kaleidoscopes, Waves and the Prepotential

Isomorphic flops are topology-changing transitions connecting two diffeomorphic families of Calabi-Yau threefolds. They correspond to the generators of certain Coxeter groups acting on the moduli space. As a consequence of these symmetries, the prepotential of 4D $\mathcal{N} = 2$ Type IIA compactifications on such varieties must assemble into Coxeter-invariant functions. We construct a database of all Coxeter symmetries from isomorphic flops in Kähler-favorable CICYs. The action of the Coxeter group on the Kähler moduli space leaves a symmetric bilinear form invariant, which we interpret as a metric and construct its associated Laplace-Beltrami operator. We argue that the Coxeter-invariant functions featured in the prepotential solve the Helmholtz equation with this Laplacian, and that the prepotential can then be resummed into a decomposition in terms of eigenfunctions of the Laplace-Beltrami operator. The convergence rate of the raw orbit sums of worldsheet instanton contributions and the resummed expressions are complementary, with the latter sharply localizing around the first few terms in the interior of the moduli space.

hep-th

Harmonic Analysis of the Instanton Prepotential

Discrete symmetries of Calabi-Yau moduli spaces, generated by isomorphic flops, constrain the instanton expansion of the 4D $\mathcal{N}=2$ Type~IIA prepotential. We show that the Coxeter-invariant functions into which the prepotential organizes are eigenfunctions of a Laplace-Beltrami operator built from the Coxeter-invariant symmetric bilinear form on the moduli space. This means that the Gromov-Witten expansion can be interpreted as a superposition of waves propagating on the Coxeter quotient of the moduli space, and its resummation is the corresponding spectral decomposition. For the dihedral Coxeter groups, separation of variables in the eigenvalue equation explains from first principles why special modified Bessel functions, ordinary Bessel functions and Jacobi theta functions appear as the natural building blocks of the prepotential, depending on whether the Coxeter rotation acts hyperbolically, elliptically, or parabolically. The resulting spectral representations converge efficiently in the interior of the moduli space, complementing the standard large-volume instanton expansion.

hep-th

Data for Mathematical Copilots: Better Ways of Presenting Proofs for Machine Learning

The datasets and benchmarks commonly used to train and evaluate the mathematical capabilities of AI-based mathematical copilots (primarily large language models) exhibit several shortcomings and misdirections. These range from a restricted scope of mathematical complexity to limited fidelity in capturing aspects beyond the final, written proof (e.g. motivating the proof, or representing the thought processes leading to a proof). These issues are compounded by a dynamic reminiscent of Goodhart's law: as benchmark performance becomes the primary target for model development, the benchmarks themselves become less reliable indicators of genuine mathematical capability. We systematically explore these limitations and contend that enhancing the capabilities of large language models, or any forthcoming advancements in AI-based mathematical assistants (copilots or ``thought partners''), necessitates a course correction both in the design of mathematical datasets and the evaluation criteria of the models' mathematical ability. In particular, it is necessary for benchmarks to move beyond the existing result-based datasets that map theorem statements directly to proofs, and instead focus on datasets that translate the richer facets of mathematical research practice into data that LLMs can learn from. This includes benchmarks that supervise the proving process and the proof discovery process itself, and we advocate for mathematical dataset developers to consider the concept of "motivated proof", introduced by G. Pólya in 1949, which can serve as a blueprint for datasets that offer a better proof learning signal, alleviating some of the mentioned limitations.

cs.LG

Fermions and Supersymmetry in Neural Network Field Theories

We introduce fermionic neural network field theories via Grassmann-valued neural networks. Free theories are obtained by a generalization of the Central Limit Theorem to Grassmann variables. This enables the realization of the free Dirac spinor at infinite width and a four fermion interaction at finite width. Yukawa couplings are introduced by breaking the statistical independence of the output weights for the fermionic and bosonic fields. A large class of interacting supersymmetric quantum mechanics and field theory models are introduced by super-affine transformations on the input that realize a superspace formalism.

hep-th

Searching for ribbons with machine learning

We apply Bayesian optimization and reinforcement learning to a problem in topology: the question of when a knot bounds a ribbon disk. This question is relevant in an approach to disproving the four-dimensional smooth Poincaré conjecture; using our programs, we rule out many potential counterexamples to the conjecture. We also show that the programs are successful in detecting many ribbon knots in the range of up to 70 crossings.

math.GT

Improving Generative Inverse Design of Rectangular Patch Antennas with Test Time Optimization

We propose a two-stage deep learning framework for the inverse design of rectangular patch antennas. Our approach leverages generative modeling to learn a latent representation of antenna frequency response curves and conditions a subsequent generative model on these responses to produce feasible antenna geometries. We further demonstrate that leveraging search and optimization techniques at test-time improves the accuracy of the generated designs and enables consideration of auxiliary objectives such as manufacturability. Our approach generalizes naturally to different design criteria, and can be easily adapted to more complex geometric design spaces.

eess.SP

Symbolic Regression with Multimodal Large Language Models and Kolmogorov Arnold Networks

We present a novel approach to symbolic regression using vision-capable large language models (LLMs) and the ideas behind Google DeepMind's Funsearch. The LLM is given a plot of a univariate function and tasked with proposing an ansatz for that function. The free parameters of the ansatz are fitted using standard numerical optimisers, and a collection of such ansätze make up the population of a genetic algorithm. Unlike other symbolic regression techniques, our method does not require the specification of a set of functions to be used in regression, but with appropriate prompt engineering, we can arbitrarily condition the generative step. By using Kolmogorov Arnold Networks (KANs), we demonstrate that ``univariate is all you need'' for symbolic regression, and extend this method to multivariate functions by learning the univariate function on each edge of a trained KAN. The combined expression is then simplified by further processing with a language model.

cs.LG

Learning Topological Invariance

Two geometric spaces are in the same topological class if they are related by certain geometric deformations. We propose machine learning methods that automate learning of topological invariance and apply it in the context of knot theory, where two knots are equivalent if they are related by ambient space isotopy. Specifically, given only the knot and no information about its topological invariants, we employ contrastive and generative machine learning techniques to map different representatives of the same knot class to the same point in an embedding vector space. An auto-regressive decoder Transformer network can then generate new representatives from the same knot class. We also describe a student-teacher setup that we use to interpret which known knot invariants are learned by the neural networks to compute the embeddings, and observe a strong correlation with the Goeritz matrix in all setups that we tested. We also develop an approach to resolving the Jones Unknot Conjecture by exploring the vicinity of the embedding space of the Jones polynomial near the locus where the unknots cluster, which we use to generate braid words with simple Jones polynomials.

math.GT

Interpretable Machine Learning for Kronecker Coefficients

We analyze the saliency of neural networks and employ interpretable machine learning models to predict whether the Kronecker coefficients of the symmetric group are zero or not. Our models use triples of partitions as input features, as well as b-loadings derived from the principal component of an embedding that captures the differences between partitions. Across all approaches, we achieve an accuracy of approximately 83% and derive explicit formulas for a decision function in terms of b-loadings. Additionally, we develop transformer-based models for prediction, achieving the highest reported accuracy of over 99%.

cs.LG

On the Learnability of Knot Invariants: Representation, Predictability, and Neural Similarity

We analyze different aspects of neural network predictions of knot invariants. First, we investigate the impact of different knot representations on the prediction of invariants and find that braid representations work in general the best. Second, we study which knot invariants are easy to learn, with invariants derived from hyperbolic geometry and knot diagrams being very easy to learn, while invariants derived from topological or homological data are harder. Predicting the Arf invariant could not be learned for any representation. Third, we propose a cosine similarity score based on gradient saliency vectors, and a joint misclassification score to uncover similarities in neural networks trained to predict related topological invariants.

math.GT

KAN: Kolmogorov-Arnold Networks

Inspired by the Kolmogorov-Arnold representation theorem, we propose Kolmogorov-Arnold Networks (KANs) as promising alternatives to Multi-Layer Perceptrons (MLPs). While MLPs have fixed activation functions on nodes ("neurons"), KANs have learnable activation functions on edges ("weights"). KANs have no linear weights at all -- every weight parameter is replaced by a univariate function parametrized as a spline. We show that this seemingly simple change makes KANs outperform MLPs in terms of accuracy and interpretability. For accuracy, much smaller KANs can achieve comparable or better accuracy than much larger MLPs in data fitting and PDE solving. Theoretically and empirically, KANs possess faster neural scaling laws than MLPs. For interpretability, KANs can be intuitively visualized and can easily interact with human users. Through two examples in mathematics and physics, KANs are shown to be useful collaborators helping scientists (re)discover mathematical and physical laws. In summary, KANs are promising alternatives for MLPs, opening opportunities for further improving today's deep learning models which rely heavily on MLPs.

cs.LG

A Heterotic Kähler Gravity and the Distance Conjecture

Deformations of the heterotic superpotential give rise to a topological holomorphic theory with similarities to both Kodaira-Spencer gravity and holomorphic Chern-Simons theory. Although the action is cubic, it is only quadratic in the complex structure deformations (the Beltrami differential). Treated separately, for large fluxes, or alternatively at large distances in the background complex structure moduli space, these fields can be integrated out to obtain a new field theory in the remaining fields, which describe the complexified hermitian and gauge degrees of freedom. We investigate properties of this new holomorphic theory, and in particular connections to the swampland distance conjecture in the context of heterotic string theory. In the process, we define a new type of symplectic cohomology theory, where the background complex structure Beltrami differential plays the role of the symplectic form.

hep-th

A Twist on Heterotic Little String Duality

In this work, we significantly expand the web of T-dualities among heterotic NS5-brane theories with eight supercharges. This is achieved by introducing twists involving outer automorphisms of discrete gauge/flavor factors and tensor multiplet permutations along the compactification circle. We assemble field theory data that we propose as invariants across T-dual theories, comprised of twisted Coulomb branch dimensions, higher group structures and flavor symmetry ranks. Using this data, we establish a detailed field theory correspondence between singularities of the compactification space, the number five-branes in the theory, and the flavor symmetry factors. The twisted theories are realized via M-theory compactifications on non-compact genus-one fibered Calabi-Yau threefolds without section. This approach allows us to prove duality of twisted and (un-)twisted theories by leveraging M/F-theory duality and identifying inequivalent torus fibrations in the same geometry. We construct several new 5D theories, including a novel type of CHL-like twisted theory where the two M9 branes are identified. Using their field theory invariants, we also construct their dual theories.

hep-th

Metric Flows with Neural Networks

We develop a general theory of flows in the space of Riemannian metrics induced by neural network gradient descent. This is motivated in part by recent advances in approximating Calabi-Yau metrics with neural networks and is enabled by recent advances in understanding flows in the space of neural networks. We derive the corresponding metric flow equations, which are governed by a metric neural tangent kernel, a complicated, non-local object that evolves in time. However, many architectures admit an infinite-width limit in which the kernel becomes fixed and the dynamics simplify. Additional assumptions can induce locality in the flow, which allows for the realization of Perelman's formulation of Ricci flow that was used to resolve the 3d Poincaré conjecture. We demonstrate that such fixed kernel regimes lead to poor learning of numerical Calabi-Yau metrics, as is expected since the associated neural networks do not learn features. Conversely, we demonstrate that well-learned numerical metrics at finite-width exhibit an evolving metric-NTK, associated with feature learning. Our theory of neural network metric flows therefore explains why neural networks are better at learning Calabi-Yau metrics than fixed kernel methods, such as the Ricci flow.

hep-th

Attractors, Geodesics, and the Geometry of Moduli Spaces

We connect recent conjectures and observations pertaining to geodesics, attractor flows, Laplacian eigenvalues and the geometry of moduli spaces by using that attractor flows are geodesics. For toroidal compactifications, attractor points are related to (degenerate) masses of the Laplacian on the target space, and also to the Laplacian on the moduli space. We also explore compactifications of M-Theory to $5$D on a Calabi-Yau threefold and argue that geodesics are unique in a special set of classes, providing further evidence for a recent conjecture by Raman and Vafa. Finally, we describe the role of the marked moduli space in $4$d $\mathcal{N} = 2$ compactifications. We study split attractor flows in an explicit example of the one-parameter family of quintics and discuss setups where flops to isomorphic Calabi-Yau manifolds exist.

hep-th

On classical de Sitter solutions and parametric control

Finding string backgrounds with de Sitter spacetime, where all approximations and corrections are controlled, is an open problem. We revisit the search for de Sitter solutions in the classical regime for specific type IIB supergravity compactifications on group manifolds, an under-explored corner of the landscape that offers an interesting testing ground for swampland conjectures. While the supergravity de Sitter solutions we obtain numerically are ambiguous in terms of their classicality, we find an analytic scaling that makes four out of six compactification radii, as well as the overall volume, arbitrarily large. This potentially provides parametric control over corrections. If we could show that these solutions, or others to be found, are fully classical, they would constitute a counterexample to conjectures stating that asymptotic de Sitter solutions do not exist. We discuss this point in great detail.

hep-th