SearcharxivSearch

arXiv subjects

Michael H. Freedman

Publications and source records attributed to Michael H. Freedman.

At least 19 recordsLinked to original sources

Artificial Intelligence and the Structure of Mathematics

Recent progress in artificial intelligence (AI) is unlocking transformative capabilities for mathematics. There is great hope that AI will help solve major open problems and autonomously discover new mathematical concepts. In this essay, we further consider how AI may open a grand perspective on mathematics by forging a new route, complementary to mathematical\textbf{ logic,} to understanding the global structure of formal \textbf{proof}\textbf{s}. We begin by providing a sketch of the formal structure of mathematics in terms of universal proof and structural hypergraphs and discuss questions this raises about the foundational structure of mathematics. We then outline the main ingredients and provide a set of criteria to be satisfied for AI models capable of automated mathematical discovery. As we send AI agents to traverse Platonic mathematical worlds, we expect they will teach us about the nature of mathematics: both as a whole, and the small ribbons conducive to human understanding. Perhaps they will shed light on the old question: "Is mathematics discovered or invented?" Can we grok the terrain of these \textbf{Platonic worlds}?

cs.AI

Correction to: Curvature growth of some 4-dimensional gradient Ricci soliton singularity models

This note corrects an error in the proof of Proposition 13 in arXiv:1903.09181 and simultaneously establishes a more general result. We prove that if $M $ is a compact connected oriented $4$-manifold with connected boundary $\partial M$, and if an unbounded number of disjoint copies of $M$ embed topologically and locally flatly in the interior of a compact $4$-manifold $N,$ then $\operatorname{Tor}H_1(\partial M;\mathbb{Z})$ is a direct double, i.e., $\operatorname{Tor}H_1(\partial M;\mathbb{Z})\cong A \oplus A$, with the linking pairing vanishing identically on the first summand, i.e., the linking pairing is split metabolic. This partially generalizes Hantzsche's theorem stating that the linking pairing for a closed $3$-manifold that embeds in $S^4$ is hyperbolic.

math.DG

Compression is all you need: Modeling Mathematics

Human mathematics (HM), the mathematics humans discover and value, is a vanishingly small subset of formal mathematics (FM), the totality of all valid deductions. We argue that HM is distinguished by its compressibility through hierarchically nested definitions, lemmas, and theorems. We model this with monoids. A mathematical deduction is a string of primitive symbols; a definition or theorem is a named substring or macro whose use compresses the string. In the free abelian monoid $A_n$, a logarithmically sparse macro set achieves exponential expansion of expressivity. In the free non-abelian monoid $F_n$, even a polynomially-dense macro set only yields linear expansion; superlinear expansion requires near-maximal density. We test these models against MathLib, a large Lean~4 library of mathematics that we take as a proxy for HM. Each element has a depth (layers of definitional nesting), a wrapped length (tokens in its definition), and an unwrapped length (primitive symbols after fully expanding all references). We find unwrapped length grows exponentially with both depth and wrapped length; wrapped length is approximately constant across all depths. These results are consistent with $A_n$ and inconsistent with $F_n$, supporting the thesis that HM occupies a polynomially-growing subset of the exponentially growing space FM. We discuss how compression, measured on the MathLib dependency graph, and a PageRank-style analysis of that graph can quantify mathematical interest and help direct automated reasoning toward the compressible regions where human mathematics lives.

cs.AI

Kolmogorov--Arnold stability

Regarding the representation theorem of Kolmogorov and Arnold (KA) as an algorithm for representing or < > functions, we test its robustness by analyzing its stability to withstand re-parameterizations of the hidden space. One may think of such re-parameterizations as the work of an adversary attempting to foil the construction of the KA outer function. We find KA to be stable under countable collections of continuous re-parameterizations, but unearth a question about the equi-continuity of the outer functions that, so far, obstructs taking limits and defeating continuous groups of re-parameterizations. This question on the regularity of the outer functions is relevant to the debate over the applicability of KA to the general theory of NNs.

cs.LG

Scale-Agnostic Kolmogorov-Arnold Geometry in Neural Networks

Recent work by Freedman and Mulligan demonstrated that shallow multilayer perceptrons spontaneously develop Kolmogorov-Arnold geometric (KAG) structure during training on synthetic three-dimensional tasks. However, it remained unclear whether this phenomenon persists in realistic high-dimensional settings and what spatial properties this geometry exhibits. We extend KAG analysis to MNIST digit classification (784 dimensions) using 2-layer MLPs with systematic spatial analysis at multiple scales. We find that KAG emerges during training and appears consistently across spatial scales, from local 7-pixel neighborhoods to the full 28x28 image. This scale-agnostic property holds across different training procedures: both standard training and training with spatial augmentation produce the same qualitative pattern. These findings reveal that neural networks spontaneously develop organized, scale-invariant geometric structure during learning on realistic high-dimensional data.

cs.LG

Spontaneous Kolmogorov-Arnold Geometry in Shallow MLPs

The Kolmogorov-Arnold (KA) representation theorem constructs universal, but highly non-smooth inner functions (the first layer map) in a single (non-linear) hidden layer neural network. Such universal functions have a distinctive local geometry, a "texture," which can be characterized by the inner function's Jacobian $J({\mathbf{x}})$, as $\mathbf{x}$ varies over the data. It is natural to ask if this distinctive KA geometry emerges through conventional neural network optimization. We find that indeed KA geometry often is produced when training vanilla single hidden layer neural networks. We quantify KA geometry through the statistical properties of the exterior powers of $J(\mathbf{x})$: number of zero rows and various observables for the minor statistics of $J(\mathbf{x})$, which measure the scale and axis alignment of $J(\mathbf{x})$. This leads to a rough understanding for where KA geometry occurs in the space of function complexity and model hyperparameters. The motivation is first to understand how neural networks organically learn to prepare input data for later downstream processing and, second, to learn enough about the emergence of KA geometry to accelerate learning through a timely intervention in network hyperparameters. This research is the "flip side" of KA-Networks (KANs). We do not engineer KA into the neural network, but rather watch KA emerge in shallow MLPs.

cs.LG

How Many Links Fit in a Box?

In an earlier note [arXiv:2301.00295] it was shown that there is an upper bound to the number of disjoint Hopf links (and certain related links) that can be embedded in the unit cube where there is a fixed separation required between the components within each copy of the Hopf link. The arguments relied on multi-linear properties of linking number and certain other link invariants. Here we produce a very similar upper bound for all non-trivial links by a more-general, entirely geometric, argument (but one which, unlike the original, has no analog in higher dimensions). Shortly after the initial paper, [arXiv:2308.08064] proved lower bounds which still provide a converse to our Theorem 1 in the case that only a bounded number of link types appear among the set $\{L_i\}$ as $N$ increases.

math.GT

Quantum Detection of Recurrent Dynamics

Quantum dynamics that explore an unexpectedly small fraction of Hilbert space is inherently interesting. Integrable systems, quantum scars, MBL, hidden tensor structures, and systems with gauge symmetries are examples. Beyond dimension and volume, spectral features such as an $O(1)$-density of periodic eigenvalues, or other spectral features, can also imply observable recurrence. Low volume dynamics will recur near its initial state $| ψ_0\rangle$ more rapidly, i.e. $\lVert\mathrm{U}^k | ψ_0\rangle - | ψ_0\rangle \rVert < ε$, is more likely to occur for modest values of $k$, when the (forward) orbit $\operatorname{closure}(\{\mathrm{U}^k\}_{k=1,2,\dots})$ is of relatively low dimension $d$ and relatively small $d$-volume. We describe simple quantum algorithms to detect such approximate recurrence. Applications include detection of certain cases of hidden tensor factorizations $\mathrm{U} \cong V^\dagger(\mathrm{U}_1\otimes \cdots \otimes \mathrm{U}_n)V$. "Hidden" refers to an unknown conjugation, e.g. $\mathrm{U}_1 \otimes \cdots \otimes \mathrm{U}_v \rightarrow V^\dagger(\mathrm{U}_1 \otimes \cdots \otimes \mathrm{U}_n)V$, which will obscure the low-volume nature of the dynamics. Hidden tensor structures have been observed to emerge both in a high energy context of operator-level spontaneous symmetry breaking [FSZ21a, FSZ21b, FSZ21c, SZBF23], and at the opposite end of the intellectual world in linguistics [Smo09, MLDS19]. We collect some observations on the computational difficulty of locating these structures and detecting related spectral information. A technical result, Appendix A, is that the language describing unitary circuits with no spectral gap (NUSG) around 1 is QMA-complete. Appendix B connects the Kolmogorov-Arnold representation theorem to hidden tensor structures.

quant-ph

The Proof of Kolmogorov-Arnold May Illuminate Neural Network Learning

Kolmogorov and Arnold, in answering Hilbert's 13th problem (in the context of continuous functions), laid the foundations for the modern theory of Neural Networks (NNs). Their proof divides the representation of a multivariate function into two steps: The first (non-linear) inter-layer map gives a universal embedding of the data manifold into a single hidden layer whose image is patterned in such a way that a subsequent dynamic can then be defined to solve for the second inter-layer map. I interpret this pattern as "minor concentration" of the almost everywhere defined Jacobians of the interlayer map. Minor concentration amounts to sparsity for higher exterior powers of the Jacobians. We present a conceptual argument for how such sparsity may set the stage for the emergence of successively higher order concepts in today's deep NNs and suggest two classes of experiments to test this hypothesis.

math.NA

Enhanced Hantzsche Theorem

A closed 3-manifold $M$ may be described up to some indeterminacy by a Heegaard diagram $\mathcal{D}$. The question "Does $M$ smoothly embed in $\mathbb{R}^4$?'' is equivalent to a property of $\mathcal{D}$ which we call $\textit{doubly unlinked}$ (DU). This perspective leads to an enhancement of Hantzsche's embedding obstruction.

math.GT

Packing Meets Topology

This note initiates an investigation of packing links into a region of Euclidean space to achieve a maximal density subject to geometric constraints. The upper bounds obtained apply only to the class of homotopically essential links and even there seem extravagantly large, leaving much working room for the interested reader.

math.GT

Universality in long-distance geometry and quantum complexity

In physics, two systems that radically differ at short scales can exhibit strikingly similar macroscopic behaviour: they are part of the same long-distance universality class. Here we apply this viewpoint to geometry and initiate a program of classifying homogeneous metrics on group manifolds by their long-distance properties. We show that many metrics on low-dimensional Lie groups have markedly different short-distance properties but nearly identical distance functions at long distances, and provide evidence that this phenomenon is even more robust in high dimensions. An application of these ideas of particular interest to physics and computer science is complexity geometry--the study of quantum computational complexity using Riemannian geometry. We argue for the existence of a large universality class of definitions of quantum complexity, each linearly related to the other, a much finer-grained equivalence than typically considered. We conjecture that a new effective metric emerges at larger complexities that describes a broad class of complexity geometries, insensitive to various choices of microscopic penalty factors. We discuss the implications for recent conjectures in quantum gravity.

hep-th

An Improved Long-Time Bishop-Gromov Theorem Using Shear

The Bishop-Gromov theorem is a comparison theorem of differential geometry that upperbounds the growth of volume of a geodesic ball in a curved space. For many spaces, this bound is far from tight. We identify a major reason the bound fails to be tight: it neglects the effect of shear. By using higher curvature invariants to lowerbound the average shear, we are able to place tighter-than-Bishop-Gromov upperbounds on the late-time growth rates of geodesic balls in homogeneous spaces with non-positive sectional curvature. We also provide concrete guidance on how our theorem can be generalized to inhomogeneous spaces, to spaces with positive sectional curvatures, and to intermediate and short times. In arXiv:2209.09288 we discovered an enhancement to the BG theorem that was strongest at early times, and that relied upon additive properties of families of Jacobi equations; in this paper we find a different enhancement at late times that connects to multiplicative properties of families of Jacobi equations. A novel feature shared by both papers is the consideration of families of equations that are not coupled but whose coefficients are correlated.

math.DG

Enhanced Bishop-Gromov Theorem

The Bishop-Gromov theorem upperbounds the rate of growth of volume of geodesic balls in a space, in terms of the most negative component of the Ricci curvature. In this paper we prove a strengthening of the Bishop-Gromov bound for homogeneous spaces. Unlike the original Bishop-Gromov bound, our enhanced bound depends not only on the most negative component of the Ricci curvature, but on the full spectrum. As a further result, for finite-volume inhomogeneous spaces, we prove an upperbound on the average rate of growth of geodesics, averaged over all starting points; this bound is stronger than the one that follows from the Bishop-Gromov theorem. Our proof makes use of the Raychaudhuri equation, of the fact that geodesic flow conserves phase-space volume, and also of a tool we introduce for studying families of correlated Jacobi equations that we call "coefficient shuffling".

math.DG

Prospecting a Possible Quadratic Wormhole Between Quantum Mechanics and Plurality

We illustrate some formal symmetries between Quadratic Funding (Buterin et al., 2019), a mechanism for the (approximately optimal) determination of public good funding levels, and the Born (1926) rule in Quantum Mechanics, which converts the wave representation into a probability distribution, through a bridging formulation we call "Quantum Quartic Finance". We suggest further directions for investigating the practical utility of these symmetries. We discuss potential interpretations in greater depth in a companion blog post.

econ.TH

Spinoza, Leibniz, Kant, and Weyl

The paper describes a funding mechanism called Quadratic Finance (QF) and deploys a bit of calculus to show that within a very clean and simple linear model QF maximizes social utility. They differentiate the social utility function. The mathematical content of this note is that by taking one further derivative, one may also deduce that QF is the unique solution.

econ.TH

Symmetry Protected Quantum Computation

We consider a model of quantum computation using qubits where it is possible to measure whether a given pair are in a singlet (total spin $0$) or triplet (total spin $1$) state. The physical motivation is that we can do these measurements in a way that is protected against revealing other information so long as all terms in the Hamiltonian are $SU(2)$-invariant. We conjecture that this model is equivalent to BQP. Towards this goal, we show: (1) this model is capable of universal quantum computation with polylogarithmic overhead if it is supplemented by single qubit $X$ and $Z$ gates. (2) Without any additional gates, it is at least as powerful as the weak model of "permutational quantum computation" of Jordan [14, 18]. (3) With postselection, the model is equivalent to PostBQP.

quant-ph

Embedding Heegaard Decompositions

A smooth embedding of a closed $3$-manifold $M$ in $\mathbb{R}^4$ may generically be composed with projection to the fourth coordinate to determine a Morse function on $M$ and hence a Heegaard splitting $M=X\cup_ΣY$. However, starting with a Heegaard splitting, we find an obstruction coming from the geometry of the curve complex $C(Σ)$ to realizing a corresponding embedding $M\hookrightarrow \mathbb{R}^4$.

math.GT