SearcharxivSearch

arXiv subjects

Jonathan P. Keating

Publications and source records attributed to Jonathan P. Keating.

At least 19 recordsLinked to original sources

When Stronger Triggers Backfire: A High-Dimensional Theory of Backdoor Attacks

Backdoor poisoning attacks behave counter-intuitively in high dimensions: stronger training triggers can help the defender. We study regularised generalised linear models on Gaussian-mixture data in the proportional regime ($p/n \to \kappa$), varying the training trigger strength $\alpha$ against a fixed test trigger. Three phenomena emerge: (i) clean test accuracy increases with $\alpha$; (ii) attack success peaks at a finite $\alpha$ and then declines; and (iii) the most damaging trigger direction is the minimum eigenvector of the data covariance. We prove all three results in closed form for the squared loss, and extend (i) and (ii) to general convex GLM losses via a Gaussian-proxy fixed-point system. We identify a finite-sample noise floor proportional to $\kappa$ as the mechanism behind (i), invisible to classical $n \gg p$ analysis. Experiments on CIFAR-10 and Gaussian surrogates match the theory closely; ResNet-18 experiments show the same phenomena beyond the convex setting.

cs.LG

The Implicit Bias of Depth: From Neural Collapse to Softmax Codes

Neural collapse (NC) describes the structured geometry that emerges in the features and weights of trained classifiers. Recent theory suggests NC can be suboptimal in deep architectures, attributing this to an explicit low-rank bias from L2 regularization. We study the deep unconstrained feature model (UFM)-equivalent to a deep linear network with orthogonal inputs-trained without regularization, to isolate how gradient descent and depth alone shape NC. We show that depth induces an implicit low-rank bias: low-rank matrices propagate norm more efficiently through successive multiplications, promoting low-rank alternatives to NC. These alternatives, we argue, correspond to softmax codes: max-margin solutions previously found in width-bottlenecked networks. Analyzing training dynamics under spectral initialization, we identify an early-time repulsion among singular values that drives low-rank emergence, and characterize how depth shrinks NC's basin of attraction. Finally, we show that some effects act in the opposite direction: for randomly initialized networks, increasing width biases training toward higher-rank solutions. Our results provide the first asymptotic and dynamic characterization of implicit bias in deep UFMs trained with unregularized multiclass cross-entropy.

cs.LG

Diagonalizing the Softmax: Hadamard Initialization for Tractable Cross-Entropy Dynamics

Cross-entropy (CE) training loss dominates deep learning practice, yet existing theory often relies on simplifications, either replacing it with squared loss or restricting to convex models, that miss essential behavior. CE and squared loss generate fundamentally different dynamics, and convex linear models cannot capture the complexities of non-convex optimization. We provide an in-depth characterization of multi-class CE optimization dynamics beyond the convex regime by analyzing a canonical two-layer linear neural network with standard-basis vectors as inputs: the simplest non-convex extension for which the implicit bias remained unknown. This model coincides with the unconstrained features model used to study neural collapse, making our work the first to prove that gradient flow on CE converges to the neural collapse geometry. We construct an explicit Lyapunov function that establishes global convergence, despite the presence of spurious critical points in the non-convex landscape. A key insight underlying our analysis is an inconspicuous finding: Hadamard Initialization diagonalizes the softmax operator, freezing the singular vectors of the weight matrices and reducing the dynamics entirely to their singular values. This technique opens a pathway for analyzing CE training dynamics well beyond our specific setting considered here.

cs.LG

Joint Moments of Characteristic Polynomials from the Orthogonal and Unitary Symplectic Groups

We establish asymptotic formulae for general joint moments of characteristic polynomials and their higher-order derivatives associated with matrices drawn randomly from the groups $\mathrm{USp}(2N)$ and $\mathrm{SO}(2N)$ in the limit as $N\to\infty$. This relates the leading-order asymptotic contribution in each case to averages over the Laguerre ensemble of random matrices. We uncover an exact connection between these joint moments and a solution of the $\sigma$-Painlev\'{e} V equation, valid for finite matrix size, as well as a connection between the leading-order asymptotic term and a solution of the $\sigma$-Painlev\'{e} III$'$ equation in the limit as $N \rightarrow \infty$. These connections enable us to derive exact formulae for joint moments for finite matrix size and for the joint moments of certain random variables arising from the Bessel point process in a recursive way. As an application, we provide a positive answer to a question proposed by Altu\u{g} et al.

math-ph

The Persistence of Neural Collapse Despite Low-Rank Bias

Neural collapse (NC) and its multi-layer variant, deep neural collapse (DNC), describe a structured geometry that occurs in the features and weights of trained deep networks. Recent theoretical work by Sukenik et al. using a deep unconstrained feature model (UFM) suggests that DNC is suboptimal under mean squared error (MSE) loss. They heuristically argue that this is due to low-rank bias induced by L2 regularization. In this work, we extend this result to deep UFMs trained with cross-entropy loss, showing that high-rank structures, including DNC, are not generally optimal. We characterize the associated low-rank bias, proving a fixed bound on the number of non-negligible singular values at global minima as network depth increases. We further analyze the loss surface, demonstrating that DNC is more prevalent in the landscape than other critical configurations, which we argue explains its frequent empirical appearance. Our results are validated through experiments in deep UFMs and deep neural networks.

cs.LG

Exchangeable arrays and integrable systems for characteristic polynomials of random matrices

The joint moments of the derivatives of the characteristic polynomial of a random unitary matrix, and also a variant of the characteristic polynomial that is real on the unit circle, in the large matrix size limit, have been studied intensively in the past twenty five years, partly in relation to conjectural connections to the Riemann zeta-function and Hardy's function. We completely settle the most general version of the problem of convergence of these joint moments, after they are suitably rescaled, for an arbitrary number of derivatives and with arbitrary positive real exponents. Our approach relies on a hidden, higher-order exchangeable structure, that of an exchangeable array. Using these probabilistic techniques, we then give a combinatorial formula for the leading order coefficient in the asymptotics of the joint moments, when the power on the characteristic polynomial itself is a positive real number and the exponents of the derivatives are integers, in terms of a finite number of finite-dimensional integrals which are explicitly computable. Finally, we develop a method, based on a class of Hankel determinants shifted by partitions, that allows us to give an exact representation of all these joint moments, for finite matrix size, in terms of derivatives of Painlev\'e V transcendents, and then for the leading order coefficient in the large-matrix limit in terms of derivatives of solutions of the $\sigma$-Painlev\'e III' equation. Equivalently, we can represent all the joint moments of power sum linear statistics of a certain determinantal point process behind this problem in terms of derivatives of $\sigma$-Painlev\'e III' transcendents. This gives an efficient way to compute all these quantities explicitly. Our methods can be used to obtain analogous results for a number of other models sharing the same features.

math.PR

Quantum interpretation of lattice paths

In the 1980s, Viennot developed a combinatorial approach to studying mixed moments of orthogonal polynomials using Motzkin paths. Recently, an alternative combinatorial model for these mixed moments based on lecture hall paths was introduced in arXiv:2311.12761. For sequences of orthogonal polynomials, we establish here a bijection between the Motzin paths and the lecture hall paths via a novel symmetric lecture hall graph. We use this connection to calculate the moments of the position operator in various separable quantum systems, such as the quantum harmonic oscillator and the hydrogen atom, showing that they may be expressed as generating functions of Motzkin paths and symmetric lecture hall paths, thereby providing a quantum interpretation for these paths. Our approach can be extended to other quantum systems where the wavefunctions are expressed in terms of orthogonal polynomials.

math-ph

Unifying Low Dimensional Spectra in Deep Learning

Low dimensional structures appear ubiquitously in the eigenspectra of deep learning matrices in classification networks trained in the overparameterized regime. While theoretical advances have aimed to explain this phenomenology, they typically succeed only in capturing subsets of the full behavior or rely on assumptions that cannot hold in practice. In this work, we provide an analytic explanation for the bulk plus outlier structure of several canonical deep learning matrices, including the Hessian, gradients, and weights. We achieve this using unconstrained feature models (UFMs), a now-common tool for studying the emergence of deep neural collapse (DNC). We show that DNC is the source of these low dimensional eigenspectra, in each case, the eigenvalues and eigenvectors can be constructed from feature means, the characterizing objects of DNC. This provides a unifying analytic explanation for a wide range of spectral phenomena in deep learning and goes beyond empirical characterizations, which typically focus on eigenvalues, by providing a detailed analysis of eigenvectors. We prove that our results hold for both linear and ReLU networks and provide numerical validation in both the modeling context and standard deep-network architectures on canonical datasets.

cs.LG

Lecture hall graphs and the Askey scheme

We establish, for every family of orthogonal polynomials in the $ q $-Askey scheme and the Askey scheme, a combinatorial model for mixed moments and coefficients in terms of paths on the lecture hall graph. This generalizes the previous results of Corteel and Kim for the little $ q $-Jacobi polynomials. We build these combinatorial models by bootstrapping, beginning with polynomials at the bottom and working towards Askey--Wilson polynomials which sit at the top of the $ q $-Askey scheme. As an application of the theory, we provide the first combinatorial proof of the symmetries in the parameters of the Askey--Wilson polynomials.

math.CO

Joint moments of higher order derivatives of CUE characteristic polynomials I: asymptotic formulae

We derive explicit asymptotic formulae for the joint moments of the $n_1$-th and $n_2$-th derivatives of the characteristic polynomials of CUE random matrices for any non-negative integers $n_1, n_2$. These formulae are expressed in terms of determinants whose entries involve modified Bessel functions of the first kind. We also express them in terms of two types of combinatorial sums. Similar results are obtained for the analogue of Hardy's $Z$-function. We use these formulae to formulate general conjectures for the joint moments of the $n_1$-th and $n_2$-th derivatives of the Riemann zeta-function and of Hardy's $Z$-function. Our conjectures are supported by comparison with results obtained previously in the number theory literature.

math-ph

Joint moments of higher order derivatives of CUE characteristic polynomials II: Structures, recursive relations, and applications

In a companion paper \cite{jon-fei}, we established asymptotic formulae for the joint moments of derivatives of the characteristic polynomials of CUE random matrices. The leading order coefficients of these asymptotic formulae are expressed as partition sums of derivatives of determinants of Hankel matrices involving I-Bessel functions, with column indices shifted by Young diagrams. In this paper, we continue the study of these joint moments and establish more properties for their leading order coefficients, including structure theorems and recursive relations. We also build a connection to a solution of the $σ$-Painlevé III$'$ equation. In the process, we give recursive formulae for the Taylor coefficients of the Hankel determinants formed from I-Bessel functions that appear and find differential equations that these determinants satisfy. The approach we establish is applicable to determinants of general Hankel matrices whose columns are shifted by Young diagrams.

math-ph

Moments of Moments of the Characteristic Polynomials of Random Orthogonal and Symplectic Matrices

Using asymptotics of Toeplitz+Hankel determinants, we establish formulae for the asymptotics of the moments of the moments of the characteristic polynomials of random orthogonal and symplectic matrices, as the matrix-size tends to infinity. Our results are analogous to those that Fahs obtained for random unitary matrices in [14]. A key feature of the formulae we derive is that the phase transitions in the moments of moments are seen to depend on the symmetry group in question in a significant way.

math-ph

The Classical Compact Groups and Gaussian Multiplicative Chaos

We consider powers of the absolute value of the characteristic polynomial of Haar distributed random orthogonal or symplectic matrices, as well as powers of the exponential of its argument, as a random measure on the unit circle minus small neighborhoods around $\pm 1$. We show that for small enough powers and under suitable normalization, as the matrix size goes to infinity, these random measures converge in distribution to a Gaussian multiplicative chaos measure. Our result is analogous to one on unitary matrices previously established by Christian Webb in [31]. We thus complete the connection between the classical compact groups and Gaussian multiplicative chaos. To prove this we establish appropriate asymptotic formulae for Toeplitz and Toeplitz+Hankel determinants with merging singularities. Using a recent formula communicated to us by Claeys et al., we are able to extend our result to the whole of the unit circle.

math-ph

On the critical-subcritical moments of moments of random characteristic polynomials: a GMC perspective

We study the 'critical moments' of subcritical Gaussian multiplicative chaos (GMCs) in dimensions $d \leq 2$. In particular, we establish a fully explicit formula for the leading order asymptotics, which is closely related to large deviation results for GMCs and demonstrates a similar universality feature. We conjecture that our result correctly describes the behaviour of analogous moments of moments of random matrices, or more generally structures which are asymptotically Gaussian and log-correlated in the entire mesoscopic scale. This is verified for an integer case in the setting of circular unitary ensemble, extending and strengthening the results of Claeys et al. and Fahs to higher-order moments.

math.PR

Hierarchical Structure in the Trace Formula

Guztwiller's Trace Formula is central to the semiclassical theory of quantum energy levels and spectral statistics in classically chaotic systems. Motivated by recent developments in Random Matrix Theory and Number Theory, we elucidate a hierarchical structure in the way periodic orbits contribute to the Trace Formula that has implications for the value distribution of spectral determinants in quantum chaotic systems.

nlin.CD

Multifractal eigenfunctions for quantum star graphs

We prove that the eigenfunctions of quantum star graphs exhibit multifractal self-similar structure in certain specified circumstances. In the semiclassical regime, when the spectral parameter and the number of vertices tend to infinity, we derive an asymptotic condition for the Mellin transform of a specified function arising from the set of bond lengths which yields an asymptotic for the Renyi entropy associated with an eigenfunction. We apply this result to show that one may construct simple quantum star graphs which satisfy a multifractal scaling law. In the low frequency regime we prove multifractality by computing the Renyi entropy in terms of a zeta function associated with the set of bond lengths. In certain arithmetic cases the fractal exponent D_q satisfies a symmetry relation around q=1/4 which arises from the functional equation of the zeta function. Our results are, in some sense, analogous to the multifractal scaling law that the authors recently proved for arithmetic Seba billiards. However, unlike in that case, we do not require arithmetic conditions to be satisfied, and nor do we rely on delicate arithmetic estimates.

math-ph

Multifractal eigenfunctions for a singular quantum billiard

Whereas much work in the mathematical literature on quantum chaos has focused on phenomena such as quantum ergodicity and scarring, relatively little is known at the rigorous level about the existence of eigenfunctions whose morphology is more complex. Quantum systems whose dynamics is intermediate between certain regimes - for example, at the transition between Anderson localized and delocalized eigenfunctions, or in systems whose classical dynamics is intermediate between integrability and chaos - have been conjectured in the physics literature to have eigenfunctions exhibiting multifractal, self-similar structure. To-date, no rigorous mathematical results have been obtained about systems of this kind in the context of quantum chaos. We give here the first rigorous proof of the existence of multifractal eigenfunctions for a widely studied class of intermediate quantum systems. Specifically, we derive an analytical formula for the Renyi entropy associated with the eigenfunctions of arithmetic Seba billiards, in the semiclassical limit, as the associated eigenvalues tend to infinity. We also prove multifractality of the ground state for more general, non-arithmetic billiards and show that the fractal exponent in this regime satisfies a symmetry relation, similar to the one predicted in the physics literature, by establishing a connection with the functional equation for Epstein's zeta function.

math-ph

On the joint moments of the characteristic polynomials of random unitary matrices

We establish the asymptotics of the joint moments of the characteristic polynomial of a random unitary matrix and its derivative for general real values of the exponents, proving a conjecture made by Hughes in 2001. Moreover, we give a probabilistic representation for the leading order coefficient in the asymptotic in terms of a real-valued random variable that plays an important role in the ergodic decomposition of the Hua-Pickrell measures. This enables us to establish connections between the characteristic function of this random variable and the $σ$-Painlevé III' equation.

math.PR