Searcharxiv⌕ Search

arXiv subjects

Jeremy Bernstein

Publications and source records attributed to Jeremy Bernstein.

30 records · Page 2Linked to original sources

signSGD: Compressed Optimisation for Non-Convex Problems

Training large neural networks requires distributing learning across multiple workers, where the cost of communicating gradients can be a significant bottleneck. signSGD alleviates this problem by transmitting just the sign of each minibatch stochastic gradient. We prove that it can get the best of both worlds: compressed gradients and SGD-level convergence rate. The relative $\ell_1/\ell_2$ geometry of gradients, noise and curvature informs whether signSGD or SGD is theoretically better suited to a particular problem. On the practical side we find that the momentum counterpart of signSGD is able to match the accuracy and convergence speed of Adam on deep Imagenet models. We extend our theory to the distributed setting, where the parameter server uses majority vote to aggregate gradient signs from each worker enabling 1-bit compression of worker-server communication in both directions. Using a theorem by Gauss we prove that majority vote can achieve the same reduction in variance as full precision distributed SGD. Thus, there is great promise for sign-based optimisation schemes to achieve fast communication and fast convergence. Code to reproduce experiments is to be found at https://github.com/jxbz/signSGD .

cs.LG↗

Stochastic Activation Pruning for Robust Adversarial Defense

Neural networks are known to be vulnerable to adversarial examples. Carefully chosen perturbations to real images, while imperceptible to humans, induce misclassification and threaten the reliability of deep learning systems in the wild. To guard against adversarial examples, we take inspiration from game theory and cast the problem as a minimax zero-sum game between the adversary and the model. In general, for such games, the optimal strategy for both players requires a stochastic policy, also known as a mixed strategy. In this light, we propose Stochastic Activation Pruning (SAP), a mixed strategy for adversarial defense. SAP prunes a random subset of activations (preferentially pruning those with smaller magnitude) and scales up the survivors to compensate. We can apply SAP to pretrained networks, including adversarially trained models, without fine-tuning, providing robustness against adversarial examples. Experiments demonstrate that SAP confers robustness against attacks, increasing accuracy and preserving calibration.

cs.LG↗

A Sommerfeld Explanation

Sommerfeld shows that the Wien displacement formula implies the existence of Planck's constant.

physics.hist-ph↗

FAPP and Non-FAPP

This is a pedagogical discussion of the foundations of the quantum theory.

physics.hist-ph↗

Another Dirac

This article discusses aspects of Dirac's work that are less familiar.

physics.hist-ph↗

SWU for You and Me

The subject of uranium isotope separation by the use of gas cenrifuges is a very active one. I present the physics of this process and some of its history.

physics.hist-ph↗

A Quantum Past

This paper discusses the formulations of the past in quantum mechanics.

quant-ph↗