SearcharxivSearch

arXiv subjects

Paul Jung

Publications and source records attributed to Paul Jung.

At least 19 recordsLinked to original sources

LLM Performance on a Real, Double-Marked GCSE Benchmark

We introduce a dataset of 32,534 double-marked real student responses to GCSE mock exams (GCSEs are the UK's national exams, taken at age ~16), spanning 328 questions across five subjects and including handwritten work. We test whether off-the-shelf large language models agree with examiners as closely as the two examiners agree with each other. We find that models overwhelmingly agree well with the examiner consensus across subjects, with the top performing models agreeing more closely with examiners than examiners agree with each other. Models achieve high scores for subjective tasks like English essay marking, as well as handling complex and messy handwritten Maths paper scripts. Agreement is uniform near the examiner line, and not massively discriminated by model size, providing cost-effective automated marking solutions.

cs.CL

Over-parameterised Shallow Neural Networks with Asymmetrical Node Scaling: Global Convergence Guarantees and Feature Learning

We consider gradient-based optimisation of wide, shallow neural networks, where the output of each hidden node is scaled by a positive parameter. The scaling parameters are non-identical, differing from the classical Neural Tangent Kernel (NTK) parameterisation. We prove that for large such neural networks, with high probability, gradient flow and gradient descent converge to a global minimum and can learn features in some sense, unlike in the NTK parameterisation. We perform experiments illustrating our theoretical results and discuss the benefits of such scaling in terms of prunability and transfer learning.

stat.ML

Deep neural networks with dependent weights: Gaussian Process mixture limit, heavy tails, sparsity and compressibility

This article studies the infinite-width limit of deep feedforward neural networks whose weights are dependent, and modelled via a mixture of Gaussian distributions. Each hidden node of the network is assigned a nonnegative random variable that controls the variance of the outgoing weights of that node. We make minimal assumptions on these per-node random variables: they are iid and their sum, in each layer, converges to some finite random variable in the infinite-width limit. Under this model, we show that each layer of the infinite-width neural network can be characterised by two simple quantities: a non-negative scalar parameter and a L\'evy measure on the positive reals. If the scalar parameters are strictly positive and the L\'evy measures are trivial at all hidden layers, then one recovers the classical Gaussian process (GP) limit, obtained with iid Gaussian weights. More interestingly, if the L\'evy measure of at least one layer is non-trivial, we obtain a mixture of Gaussian processes (MoGP) in the large-width limit. The behaviour of the neural network in this regime is very different from the GP regime. One obtains correlated outputs, with non-Gaussian distributions, possibly with heavy tails. Additionally, we show that, in this regime, the weights are compressible, and some nodes have asymptotically non-negligible contributions, therefore representing important hidden features. Many sparsity-promoting neural network models can be recast as special cases of our approach, and we discuss their infinite-width limits; we also present an asymptotic analysis of the pruning error. We illustrate some of the benefits of the MoGP regime over the GP regime in terms of representation learning and compressibility on simulated, MNIST and Fashion MNIST datasets.

stat.ML

Large deviations in the quantum quasi-1D jellium

Wigner's jellium is a model for a gas of electrons. The model consists of $N$ unit negatively charged particles lying in a sea of neutralizing homogeneous positive charge spread out according to Lebesgue measure, and interactions are governed by the Coulomb potential. In this work we consider the quantum jellium on quasi-one-dimensional spaces with Maxwell-Boltzmann statistics. Using the Feynman-Kac representation, we replace particle locations with Brownian bridges. We then adapt the approach of Leblé and Serfaty (2017) to prove a process-level large deviation principle for the empirical fields of the Brownian bridges.

math.PR

$α$-Stable convergence of heavy-tailed infinitely-wide neural networks

We consider infinitely-wide multi-layer perceptrons (MLPs) which are limits of standard deep feed-forward neural networks. We assume that, for each layer, the weights of an MLP are initialized with i.i.d. samples from either a light-tailed (finite variance) or heavy-tailed distribution in the domain of attraction of a symmetric $α$-stable distribution, where $α\in(0,2]$ may depend on the layer. For the bias terms of the layer, we assume i.i.d. initializations with a symmetric $α$-stable distribution having the same $α$ parameter of that layer. We then extend a recent result of Favaro, Fortini, and Peluchetti (2020), to show that the vector of pre-activation values at all nodes of a given hidden layer converges in the limit, under a suitable scaling, to a vector of i.i.d. random variables with symmetric $α$-stable distributions.

stat.ML

At the edge of a one-dimensional jellium

We consider a one-dimensional classical Wigner jellium, not necessarily charge neutral, for which the electrons are allowed to exist beyond the support of the background charge. The model can be seen as a one-dimensional Coulomb gas in which the external field is generated by a smeared background on an interval. It is a true one-dimensional Coulomb gas and not a one-dimensional log-gas. The system exists if and only if the total background charge is greater than the number of electrons minus one. For various backgrounds, we show convergence to point processes, at the edge of the support of the background. In particular, this provides asymptotic analysis of the fluctuations of the right-most particle. Our analysis reveals that these fluctuations are not universal, in the sense that depending on the background, the tails range anywhere from exponential to Gaussian-like behavior, including for instance Tracy-Widom-like behavior. We also obtain a Renyi-type probabilistic representation for the order statistics of the particle system beyond the support of the background.

math.PR

Extreme eigenvalue statistics of $m$-dependent heavy-tailed matrices

We analyze the largest eigenvalue statistics of m-dependent heavy-tailed Wigner matrices as well as the associated sample covariance matrices having entry-wise regularly varying tail distributions with parameter $0<α<4$. Our analysis extends results in the previous literature for the corresponding random matrices with independent entries above the diagonal, by allowing for m-dependence between the entries of a given matrix. We prove that the limiting point process of extreme eigenvalues is a Poisson cluster process.

math.PR

The Life and Mathematical Legacy of Thomas M. Liggett

Thomas Milton Liggett was a world renowned UCLA probabilist, famous for his monograph Interacting Particle Systems. He passed away peacefully on May 12, 2020. This is a perspective article in memory of both Tom Liggett the person and Tom Liggett the mathematician.

math.PR

A Generalization of Hierarchical Exchangeability on Trees to Directed Acyclic Graphs

Motivated by the problem of designing inference-friendly Bayesian nonparametric models in probabilistic programming languages, we introduce a general class of partially exchangeable random arrays which generalizes the notion of hierarchical exchangeability introduced in Austin and Panchenko (2014). We say that our partially exchangeable arrays are DAG-exchangeable since their partially exchangeable structure is governed by a collection of Directed Acyclic Graphs. More specifically, such a random array is indexed by $\mathbb{N}^{|V|}$ for some DAG $G=(V,E)$, and its exchangeability structure is governed by the edge set $E$. We prove a representation theorem for such arrays which generalizes the Aldous-Hoover and Austin-Panchenko representation theorems.

math.PR

Macroscopic and edge behavior of a planar jellium

We consider a planar Coulomb gas in which the external potential is generated by a smeared uniform background of opposite-sign charge on a disc. This model can be seen as a two-dimensional Wigner jellium, not necessarily charge neutral, and with particles allowed to exist beyond the support of the smeared charge. The full space integrability condition requires low enough temperature or high enough total smeared charge. This condition does not allow at the same time, total charge neutrality and determinantal structure. The model shares similarities with both the complex Ginibre ensemble and the Forrester--Krishnapur spherical ensemble of random matrix theory. In particular, for a certain regime of temperature and total charge, the equilibrium measure is uniform on a disc as in the Ginibre ensemble, while the modulus of the farthest particle has heavy-tailed fluctuations as in the Forrester--Krishnapur spherical ensemble. We also touch on a higher temperature regime producing a crossover equilibrium measure, as well as a transition to Gumbel edge fluctuations. More results in the same spirit on edge fluctuations are explored by the second author together with Raphael Butez.

math.PR

A note on invariance of the Cauchy and related distributions

It is known that if $f$ is an analytic self map of the complex upper half-plane which also maps $\mathbb{R}\cup\{\infty\}$ to itself, and $f(i)=i$, then $f$ preserves the Cauchy distribution. This note concerns three results related to the above fact.

math.PR

Convergence to $α$-stable Lévy motion for chaotic billiards with several cusps at flat points

We consider billiards with several possibly non-isometric and asymmetric cusps at flat points; the case of a single symmetric cusp was studied previously in Zhang (2017) and Jung & Zhang (2018). In particular, we show that properly normalized Birkhoff sums of Hölder observables, with respect to the billiard map, converge in Skorokhod's $M_1$-topology to an $α$-stable Lévy motion, where $α$ depends on the `curvature' of the flattest points and the skewness parameter $ξ$ depends on the values of the observable at those same points. Previously, Jung & Zhang (2018) proved convergence of the one-point marginals to totally skewed $α$-stable distributions for a single symmetric cusp. The limits we prove here are stronger, since they are in the functional sense, but also allow for more varied behaviour due to the presence of multiple cusps. In particular, the general limits we obtain allow for any skewness parameter, as opposed to just the totally skewed cases. We also show that convergence in the stronger $J_1$-topology is not possible.

math.DS

Stable laws for chaotic billiards with cusps at flat points

We consider billiards with a single cusp where the walls meeting at the vertex of the cusp have zero one-sided curvature, thus forming a flat point at the vertex. For Hölder continuous observables, we show that properly normalized Birkoff sums, with respect to the billiard map, converge in law to a totally skewed $α$-stable law.

math-ph

On the speed and spectrum of mean-field random walks among random conductances

We study random walk among random conductance (RWRC) on complete graphs with N vertices. The conductances are i.i.d. and the sum of conductances emanating from a single vertex asymptotically has an infinitely divisible distribution corresponding to a Lévy subordinator with infinite mass at 0. We show that, under suitable conditions, the empirical spectral distribution of the random transition matrix associated to the RWRC converges weakly, as N goes to infinity, to a symmetric deterministic measure on [-1,1], in probability with respect to the randomness of the conductances. In short time scales, the limiting underlying graph of the RWRC is a Poisson Weighted Infinite Tree, and we analyze the RWRC on this limiting tree. In particular, we show that the transient RWRC exhibits a phase transition in that it has positive or weakly zero speed when the mean of the largest conductance is finite or infinite, respectively.

math.PR

Delocalization and Limiting Spectral Distribution of Erdős-Rényi Graphs with Constant Expected Degree

We consider Erdős-Rényi graphs $G(n,p_n)$ with large constant expected degree $λ$ and $p_n=λ/n$. Bordenave and Lelarge (2010) showed that the infinite-volume limit, in the Benjamini-Schramm topology, is a Galton-Watson tree with offspring distribution Pois($λ$) and the mean spectrum at the root of this tree has unbounded support and corresponds to the limiting spectral distribution of $G(n,p_n)$ as $n\to\infty$. We show that if one weights the edges by $1/\sqrtλ$ and sends $λ\to\infty$, then the support mostly vanishes and in fact, the limiting spectral distributions converge weakly to a semicircle distribution. We also find that for large $λ$, there is an orthonormal eigenvector basis of $G(n,p_n)$ such that most of the vectors delocalize with respect to the infinity norm, as $n\to\infty$. Our delocalization result provides a variant on a result of Tran, Vu and Wang (2013).

math.PR

Levy-Khintchine random matrices and the Poisson weighted infinite skeleton tree

We study a class of Hermitian random matrices which includes and generalizes Wigner matrices, heavy-tailed random matrices, and sparse random matrices such as the adjacency matrices of Erdos-Renyi random graphs with p ~ 1/N. Our NxN random matrices have real entries which are i.i.d. up to symmetry. The distributions may depend on N, however, the sums of rows must converge in distribution; it is then well-known that the limiting distributions are infinitely divisible. We show that a limiting empirical spectral distribution (LSD) exists, and via local weak convergence of associated graphs, the LSD corresponds to the spectral measure at the root of a Poisson weighted infinite skeleton tree. This graph is formed by connecting infinitely many Poisson weighted infinite trees using a backbone structure. One example covered by the results are matrices with i.i.d. entries having infinite second moments, but normalized to be in the Gaussian domain of attraction. In this case, the limiting graph is just the positive integer line rooted at 1, and as expected, the LSD is Wigner's semi-circle law. The results also extend to self-adjoint complex matrices and also to Wishart matrices.

math.PR