SearcharxivSearch

arXiv subjects

Dustin G. Mixon

Publications and source records attributed to Dustin G. Mixon.

At least 19 recordsLinked to original sources

Towards a mathematical theory of superposition

We develop a mathematical theory of superposition in neural networks using tools from frame theory and compressed sensing. In our model, a sparse binary vector \(x\) of active features is encoded through an overcomplete dictionary \(W\), and feature recovery is performed by applying \(\operatorname{ReLU}(W^\top W x+b)\) with an appropriate bias vector \(b\). We prove several recovery theorems for this model. In the random-support setting, we establish high-probability support recovery for nearly tight, low-coherence dictionaries, with guarantees when the expected sparsity is up to order \(d/\log n\). In the worst-case support setting, we give a sharp and computable criterion for which sparsity levels permit support recovery. We apply this criterion to Gaussian random matrices and equiangular tight frames. For real equiangular tight frames with \(n>d+1\), we determine the exact recovery threshold in terms of the coherence. The proof of this result for real equiangular tight frames relies on a novel characterization---which should be of independent interest to frame theorists---of the distribution of signs in the Gram matrix.

stat.ML

Asymptotically optimal approximate Hadamard matrices

An approximate Hadamard matrix is a well-conditioned square matrix with all entries in $\{\pm1\}$. We measure the quality of a matrix by its condition number, i.e., the ratio of its largest and smallest singular values. We prove that for any fixed positive $α<17/92$, every sufficiently large dimension admits an approximate Hadamard matrix with condition number at most $1+n^{-α}$. In particular, the smallest possible condition number tends to $1$ as $n\to\infty$. Conversely, there exists an absolute constant $c>0$ such that for every sufficiently large $n\not\equiv0\pmod4$, every $n\times n$ matrix with entries in $\{\pm1\}$ has condition number at least $1+c(\log n)/n$. Along the way, we resolve a problem of Jaming and Matolcsi concerning flat orthogonal matrices, and we conclude by describing several explicit infinite families of approximate Hadamard matrices.

math.CO

HRT counterexamples with exponential tails

We build on the recent breakthrough of Faulhuber, Petersen, van Velthoven, and Voigtlaender that disproved the HRT conjecture with a Schwartz function and a $12$-point configuration. We give a human-readable treatment of their mechanism and find HRT counterexample functions with exponential (or faster) decay. By a result of Bownik and Speegle, this is the fastest possible decay for an HRT counterexample, up to a logarithmic factor in the exponent.

math.FA

Short spherical $t$-design curves

We study the minimum arclength of spherical $t$-design curves, i.e., closed rectifiable curves on $S^d$ whose normalized arclength measure exactly integrates every polynomial of degree at most $t$. We prove an explicit spectral lower bound that is sharp for $t=1$ in all spheres and for $t=2$ in every odd-dimensional sphere, yielding the first exact optimality results for spherical $t$-design curves with $t>1$. For even-dimensional spheres, we construct $2$-design curves whose lengths asymptotically match the lower bound as $d\to\infty$, and in $S^2$, we use numerical optimization and the calculus of variations to derive a candidate for the shortest $2$-design curve.

math.CO

The Singer-Zauner gap for equiangular tight frames

We show that there does not exist a complex $d\times n$ equiangular tight frame with \[ d^2-d+1<n<d^2. \] The proof, which originated from an internal model at OpenAI, mimics the relationship between real equiangular tight frames and strongly regular graphs.

math.FA

Approximation theorems in bilipschitz invariant theory

Bilipschitz invariant theory concerns low-distortion embeddings of orbit spaces into Euclidean space. To date, embeddings with the smallest-possible distortion are known for only a few cases, to include: (a) planar rotations, (b) real phase retrieval, and (c) finite reflection groups. Here, we prove that for all three of these cases, the smallest possible distortion is nearly achieved by a composition of a "max filter bank" with a linear transformation. Our proof amounts to a two-step process: first, we show it suffices to demonstrate a certain inclusion of Lipschitz function spaces, and second, we prove that inclusion, using fundamentally different approaches for the three cases. We also show that these cases interact differently with a few related function spaces, which suggests that a unified treatment would be nontrivial.

math.FA

Neural collapse in the orthoplex regime

When training a neural network for classification, the feature vectors of the training set are known to collapse to the vertices of a regular simplex, provided the dimension $d$ of the feature space and the number $n$ of classes satisfies $n\leq d+1$. This phenomenon is known as neural collapse. For other applications like language models, one instead takes $n\gg d$. Here, the neural collapse phenomenon still occurs, but with different emergent geometric figures. We characterize these geometric figures in the orthoplex regime where $d+2\leq n\leq 2d$. The techniques in our analysis primarily involve Radon's theorem and convexity.

cs.LG

SUNLayer: Stable denoising with generative networks

Deep neural networks are often used to implement powerful generative models for real-world data. Notable applications include image denoising, as well as other classical inverse problems like compressed sensing and super-resolution. To provide a rigorous but simplified analysis of generative models, in this work, we introduce an elegant theoretical framework based on spherical harmonics, namely \textbf{SUNLayer}. Our theoretical framework identifies explicit conditions on activation functions that guarantee denoising under local optimization. Numerical experiments examine the theoretical properties on commonly used activation functions and demonstrate their stable denoising performance.

cs.LG

Totally symmetric Grassmannian codes

We introduce a general technique to construct tight fusion frames with prescribed symmetries. Applying this technique with a prescription for "all the symmetries", we construct a new family of equi-isoclinic tight fusion frames (EITFFs), which consequently form optimal Grassmannian codes. By virtue of their construction, our EITFFs have the remarkable property of total symmetry: any permutation of subspaces can be achieved by an appropriate unitary.

math.CO

Forbidden Sidon subsets of perfect difference sets, featuring a human-assisted proof

We resolve a $1000 Erdős prize problem, complete with formal verification generated by a large language model. In over a dozen papers, beginning in 1976 and spanning two decades, Paul Erdős repeatedly posed one of his "favourite" conjectures: every finite Sidon set can be extended to a finite perfect difference set. We establish that {1, 2, 4, 8, 13} is a counterexample to this conjecture. During the preparation of this paper, we discovered that although this problem was presumed to be open for half a century, Marshall Hall, Jr. published a different counterexample three decades before Erdős first posed the problem. With a healthy skepticism of this apparent oversight, and out of an abundance of caution, we used ChatGPT to vibe code a Lean proof of both Hall's and our counterexamples.

math.CO

Testing isomorphism between tuples of subspaces

Given two tuples of subspaces, can you tell whether the tuples are isomorphic? We develop theory and algorithms to address this fundamental question. We focus on isomorphisms in which the ambient vector space is acted on by either a unitary group or general linear group. If isomorphism also allows permutations of the subspaces, then the problem is at least as hard as graph isomorphism. Otherwise, we provide a variety of polynomial-time algorithms with Matlab implementations to test for isomorphism. Keywords: subspace isomorphism, Grassmannian, Bargmann invariants, $H^\ast$-algebras, quivers, graph isomorphism

math.MG

BalLOT: Balanced $k$-means clustering with optimal transport

We consider the fundamental problem of balanced $k$-means clustering. In particular, we introduce an optimal transport approach to alternating minimization called BalLOT, and we show that it delivers a fast and effective solution to this problem. We establish this with a variety of numerical experiments before proving several theoretical guarantees. First, we prove that for generic data, BalLOT produces integral couplings at each step. Next, we perform a landscape analysis to provide theoretical guarantees for both exact and partial recoveries of planted clusters under the stochastic ball model. Finally, we propose initialization schemes that achieve one-step recovery of planted clusters.

stat.ML

The independence and clique cover numbers of the squarefree graph

We determine the largest subset $A\subseteq \{1,\dotsc,n\}$ such that for all $a,b\in A$, the product $ab$ is not squarefree. Specifically, the maximum size is achieved by the complement of the odd squarefree numbers. This resolves a problem of Paul Erdős and András Sárközy from 1992.

math.CO

Asymmetric SICs over finite fields

Zauner's conjecture concerns the existence of $d^2$ equiangular lines in $\mathbb{C}^d$; such a system of lines is known as a SIC. In this paper, we construct infinitely many new SICs over finite fields. While all previously known SICs exhibit Weyl--Heisenberg symmetry, some of our new SICs exhibit trivial automorphism groups. We conjecture that such \textit{totally asymmetric} SICs exist in infinitely many dimensions in the finite field setting.

math.MG

Estimating the Euclidean distortion of an orbit space

Given a finite-dimensional inner product space $V$ and a group $G$ of isometries, we consider the problem of embedding the orbit space $V/G$ into a Hilbert space in a way that preserves the quotient metric as well as possible. This inquiry is motivated by applications to invariant machine learning. We introduce several new theoretical tools before using them to tackle various fundamental instances of this problem.

math.MG

On the clustering behavior of sliding windows

Things can go spectacularly wrong when clustering timeseries data that has been preprocessed with a sliding window. We highlight three surprising failures that emerge depending on how the window size compares with the timeseries length. In addition to computational examples, we present theoretical explanations for each of these failure modes.

cs.LG