Searcharxiv⌕ Search

arXiv subjects

Andrzej Banburski

Publications and source records attributed to Andrzej Banburski.

15 recordsLinked to original sources

Non-local Field Theory from Matrix Models

We show that a class of matrix theories can be understood as an extension of quantum field theory which has non-local interactions. This reformulation is based on the Wigner-Weyl transformation, and the interactions take the form of Moyal product on a doubled geometry. We recover local dynamics on the spacetime as a low-energy limit. This framework opens up the possibility for studying novel high-energy phenomena, including the unification of gauge and geometric symmetries in a gauge theory.

hep-th↗

Neural-guided, Bidirectional Program Search for Abstraction and Reasoning

One of the challenges facing artificial intelligence research today is designing systems capable of utilizing systematic reasoning to generalize to new tasks. The Abstraction and Reasoning Corpus (ARC) measures such a capability through a set of visual reasoning tasks. In this paper we report incremental progress on ARC and lay the foundations for two approaches to abstraction and reasoning not based in brute-force search. We first apply an existing program synthesis system called DreamCoder to create symbolic abstractions out of tasks solved so far, and show how it enables solving of progressively more challenging ARC tasks. Second, we design a reasoning algorithm motivated by the way humans approach ARC. Our algorithm constructs a search graph and reasons over this graph structure to discover task solutions. More specifically, we extend existing execution-guided program synthesis approaches with deductive reasoning based on function inverse semantics to enable a neural-guided bidirectional search algorithm. We demonstrate the effectiveness of the algorithm on three domains: ARC, 24-Game tasks, and a 'double-and-add' arithmetic puzzle.

cs.AI↗

Distribution of Classification Margins: Are All Data Equal?

Recent theoretical results show that gradient descent on deep neural networks under exponential loss functions locally maximizes classification margin, which is equivalent to minimizing the norm of the weight matrices under margin constraints. This property of the solution however does not fully characterize the generalization performance. We motivate theoretically and show empirically that the area under the curve of the margin distribution on the training set is in fact a good measure of generalization. We then show that, after data separation is achieved, it is possible to dynamically reduce the training set by more than 99% without significant loss of performance. Interestingly, the resulting subset of "high capacity" features is not consistent across different training runs, which is consistent with the theoretical claim that all training points should converge to the same asymptotic margin under SGD and in the presence of both batch normalization and weight decay.

cs.LG↗

Hierarchically Compositional Tasks and Deep Convolutional Networks

The main success stories of deep learning, starting with ImageNet, depend on deep convolutional networks, which on certain tasks perform significantly better than traditional shallow classifiers, such as support vector machines, and also better than deep fully connected networks; but what is so special about deep convolutional networks? Recent results in approximation theory proved an exponential advantage of deep convolutional networks with or without shared weights in approximating functions with hierarchical locality in their compositional structure. More recently, the hierarchical structure was proved to be hard to learn from data, suggesting that it is a powerful prior embedded in the architecture of the network. These mathematical results, however, do not say which real-life tasks correspond to input-output functions with hierarchical locality. To evaluate this, we consider a set of visual tasks where we disrupt the local organization of images via "deterministic scrambling" to later perform a visual task on these images structurally-altered in the same way for training and testing. For object recognition we find, as expected, that scrambling does not affect the performance of shallow or deep fully connected networks contrary to the out-performance of convolutional networks. Not all tasks involving images are however affected. Texture perception and global color estimation are much less sensitive to deterministic scrambling showing that the underlying functions corresponding to these tasks are not hierarchically local; and also counter-intuitively showing that these tasks are better approximated by networks that are not deep (texture) nor convolutional (color). Altogether, these results shed light into the importance of matching a network architecture with its embedded prior of the task to be learned.

cs.LG↗

Biologically Inspired Mechanisms for Adversarial Robustness

A convolutional neural network strongly robust to adversarial perturbations at reasonable computational and performance cost has not yet been demonstrated. The primate visual ventral stream seems to be robust to small perturbations in visual stimuli but the underlying mechanisms that give rise to this robust perception are not understood. In this work, we investigate the role of two biologically plausible mechanisms in adversarial robustness. We demonstrate that the non-uniform sampling performed by the primate retina and the presence of multiple receptive fields with a range of receptive field sizes at each eccentricity improve the robustness of neural networks to small adversarial perturbations. We verify that these two mechanisms do not suffer from gradient obfuscation and study their contribution to adversarial robustness through ablation studies.

cs.LG↗

Double descent in the condition number

In solving a system of $n$ linear equations in $d$ variables $Ax=b$, the condition number of the $n,d$ matrix $A$ measures how much errors in the data $b$ affect the solution $x$. Estimates of this type are important in many inverse problems. An example is machine learning where the key task is to estimate an underlying function from a set of measurements at random points in a high dimensional space and where low sensitivity to error in the data is a requirement for good predictive performance. Here we discuss the simple observation, which is known but surprisingly little quoted (see Theorem 4.2 in \cite{Brgisser:2013:CGN:2526261}): when the columns of $A$ are random vectors, the condition number of $A$ is highest if $d=n$, that is when the inverse of $A$ exists. An overdetermined system ($n>d$) as well as an underdetermined system ($n<d$), for which the pseudoinverse must be used instead of the inverse, typically have significantly better, that is lower, condition numbers. Thus the condition number of $A$ plotted as function of $d$ shows a double descent behavior with a peak at $d=n$.

cs.LG↗

Theory III: Dynamics and Generalization in Deep Networks

The key to generalization is controlling the complexity of the network. However, there is no obvious control of complexity -- such as an explicit regularization term -- in the training of deep networks for classification. We will show that a classical form of norm control -- but kind of hidden -- is present in deep networks trained with gradient descent techniques on exponential-type losses. In particular, gradient descent induces a dynamics of the normalized weights which converge for $t \to \infty$ to an equilibrium which corresponds to a minimum norm (or maximum margin) solution. For sufficiently large but finite $ρ$ -- and thus finite $t$ -- the dynamics converges to one of several margin maximizers, with the margin monotonically increasing towards a limit stationary point of the flow. In the usual case of stochastic gradient descent, most of the stationary points are likely to be convex minima corresponding to a constrained minimizer -- the network with normalized weights-- which corresponds to vanishing regularization. The solution has zero generalization gap, for fixed architecture, asymptotically for $N \to \infty$, where $N$ is the number of training examples. Our approach extends some of the original results of Srebro from linear networks to deep networks and provides a new perspective on the implicit bias of gradient descent. We believe that the elusive complexity control we describe is responsible for the puzzling empirical finding of good predictive performance by deep networks, despite overparametrization.

cs.LG↗

Theoretical Issues in Deep Networks: Approximation, Optimization and Generalization

While deep learning is successful in a number of applications, it is not yet well understood theoretically. A satisfactory theoretical characterization of deep learning however, is beginning to emerge. It covers the following questions: 1) representation power of deep networks 2) optimization of the empirical risk 3) generalization properties of gradient descent techniques --- why the expected error does not suffer, despite the absence of explicit regularization, when the networks are overparametrized? In this review we discuss recent advances in the three areas. In approximation theory both shallow and deep networks have been shown to approximate any continuous functions on a bounded domain at the expense of an exponential number of parameters (exponential in the dimensionality of the function). However, for a subset of compositional functions, deep networks of the convolutional type can have a linear dependence on dimensionality, unlike shallow networks. In optimization we discuss the loss landscape for the exponential loss function and show that stochastic gradient descent will find with high probability the global minima. To address the question of generalization for classification tasks, we use classical uniform convergence results to justify minimizing a surrogate exponential-type loss function under a unit norm constraint on the weight matrix at each layer -- since the interesting variables for classification are the weight directions rather than the weights. Our approach, which is supported by several independent new results, offers a solution to the puzzle about generalization performance of deep overparametrized ReLU networks, uncovering the origin of the underlying hidden complexity control.

cs.LG↗

A Surprising Linear Relationship Predicts Test Performance in Deep Networks

Given two networks with the same training loss on a dataset, when would they have drastically different test losses and errors? Better understanding of this question of generalization may improve practical applications of deep networks. In this paper we show that with cross-entropy loss it is surprisingly simple to induce significantly different generalization performances for two networks that have the same architecture, the same meta parameters and the same training error: one can either pretrain the networks with different levels of "corrupted" data or simply initialize the networks with weights of different Gaussian standard deviations. A corollary of recent theoretical results on overfitting shows that these effects are due to an intrinsic problem of measuring test performance with a cross-entropy/exponential-type loss, which can be decomposed into two components both minimized by SGD -- one of which is not related to expected classification performance. However, if we factor out this component of the loss, a linear relationship emerges between training and test losses. Under this transformation, classical generalization bounds are surprisingly tight: the empirical/training loss is very close to the expected/test loss. Furthermore, the empirical relation between classification error and normalized cross-entropy loss seem to be approximately monotonic

cs.LG↗

Theory IIIb: Generalization in Deep Networks

A main puzzle of deep neural networks (DNNs) revolves around the apparent absence of "overfitting", defined in this paper as follows: the expected error does not get worse when increasing the number of neurons or of iterations of gradient descent. This is surprising because of the large capacity demonstrated by DNNs to fit randomly labeled data and the absence of explicit regularization. Recent results by Srebro et al. provide a satisfying solution of the puzzle for linear networks used in binary classification. They prove that minimization of loss functions such as the logistic, the cross-entropy and the exp-loss yields asymptotic, "slow" convergence to the maximum margin solution for linearly separable datasets, independently of the initial conditions. Here we prove a similar result for nonlinear multilayer DNNs near zero minima of the empirical loss. The result holds for exponential-type losses but not for the square loss. In particular, we prove that the weight matrix at each layer of a deep network converges to a minimum norm solution up to a scale factor (in the separable case). Our analysis of the dynamical system corresponding to gradient descent of a multilayer network suggests a simple criterion for ranking the generalization performance of different zero minimizers of the empirical loss.

cs.LG↗

A simpler way of imposing simplicity constraints

We investigate a way of imposing simplicity constraints in a holomorphic Spin Foam model that we recently introduced. Rather than imposing the constraints on the boundary spin network, as is usually done, one can impose the constraints directly on the Spin Foam propagator. We find that the two approaches have the same leading asymptotic behaviour, with differences appearing at higher order. This allows us to obtain a model that greatly simplifies calculations, but still has Regge Calculus as its semi-classical limit.

gr-qc↗

Pachner moves in a 4d Riemannian holomorphic Spin Foam model

In this work we study a Spin Foam model for 4d Riemannian gravity, and propose a new way of imposing the simplicity constraints that uses the recently developed holomorphic representation. Using the power of the holomorphic integration techniques, and with the introduction of two new tools: the homogeneity map and the loop identity, for the first time we give the analytic expressions for the behaviour of the Spin Foam amplitudes under 4-dimensional Pachner moves. It turns out that this behaviour is controlled by an insertion of nonlocal mixing operators. In the case of the 5-1 move, the expression governing the change of the amplitude can be interpreted as a vertex renormalisation equation. We find a natural truncation scheme that allows us to get an invariance up to an overall factor for the 4-2 and 5-1 moves, but not for the 3-3 move. The study of the divergences shows that there is a range of parameter space for which the 4-2 move is finite while the 5-1 move diverges. This opens up the possibility to recover diffeomorphism invariance in the continuum limit of Spin Foam models for 4D Quantum Gravity.

gr-qc↗

Snyder Momentum Space in Relative Locality

The standard approaches of phenomenology of Quantum Gravity have usually explicitly violated Lorentz invariance, either in the dispersion relation or in the addition rule for momenta. We investigate whether it is possible in 3+1 dimensions to have a non local deformation that preserves fully Lorentz invariance, as it is the case in 2+1D Quantum Gravity. We answer positively to this question and show for the first time how to construct a homogeneously curved momentum space preserving the full action of the Lorentz group in dimension 4 and higher, despite relaxing locality. We study the property of this relative locality deformation and show that this space leads to a noncommutativity related to Snyder spacetime.

gr-qc↗

Twisting loops and global momentum non-conservation in Relative Locality

Recent work in Relative Locality has shown that the theory allows for a solution of an on-shell causal loop. We show that the theory contains a different type of a loop in which locally momenta are conserved, but there is no global momentum conservation. Thus a freely propagating particle can decay into two particles, which later recombine to give a particle with momentum and mass different than the original one.

gr-qc↗

The Production and Discovery of True Muonium in Fixed-Target Experiments

Upcoming fixed-target experiments designed to search for new sub-GeV forces will also have sensitivity to the never before observed True Muonium atom, a bound state of a muon and anti-muon. We describe the production and decay characteristics of True Muonium relevant to these experiments. Importantly, we find that secondary production mechanisms dominate over primary production for the long-lived 2S and 2P states, leading to total yields an order of magnitude larger than naive estimates previously suggested. We present yield estimates for True Muonium as a function of energy fraction and decay length, useful for guiding future experimental studies. Discovery and measurement prospects appear very favorable.

hep-ph↗