SearcharxivSearch

arXiv subjects

Han Kheng Teoh

Publications and source records attributed to Han Kheng Teoh.

3 recordsLinked to original sources

The Training Process of Many Deep Networks Explores the Same Low-Dimensional Manifold

We develop information-geometric techniques to analyze the trajectories of the predictions of deep networks during training. By examining the underlying high-dimensional probabilistic models, we reveal that the training process explores an effectively low-dimensional manifold. Networks with a wide range of architectures, sizes, trained using different optimization methods, regularization techniques, data augmentation techniques, and weight initializations lie on the same manifold in the prediction space. We study the details of this manifold to find that networks with different architectures follow distinguishable trajectories but other factors have a minimal influence; larger networks train along a similar manifold as that of smaller networks, just faster; and networks initialized at very different parts of the prediction space converge to the solution along a similar manifold.

cs.LG

A picture of the space of typical learnable tasks

We develop information geometric techniques to understand the representations learned by deep networks when they are trained on different tasks using supervised, meta-, semi-supervised and contrastive learning. We shed light on the following phenomena that relate to the structure of the space of tasks: (1) the manifold of probabilistic models trained on different tasks using different representation learning methods is effectively low-dimensional; (2) supervised learning on one task results in a surprising amount of progress even on seemingly dissimilar tasks; progress on other tasks is larger if the training task has diverse classes; (3) the structure of the space of tasks indicated by our analysis is consistent with parts of the Wordnet phylogenetic tree; (4) episodic meta-learning algorithms and supervised learning traverse different trajectories during training but they fit similar models eventually; (5) contrastive and semi-supervised learning methods traverse trajectories similar to those of supervised learning. We use classification tasks constructed from the CIFAR-10 and Imagenet datasets to study these phenomena.

cs.LG

Visualizing probabilistic models in Minkowski space with intensive symmetrized Kullback-Leibler embedding

We show that the predicted probability distributions for any $N$-parameter statistical model taking the form of an exponential family can be explicitly and analytically embedded isometrically in a $N{+}N$-dimensional Minkowski space. That is, the model predictions can be visualized as control parameters are varied, preserving the natural distance between probability distributions. All pairwise distances between model instances are given by the symmetrized Kullback-Leibler divergence. We give formulas for these intensive symmetrized Kullback Leibler (isKL) coordinate embeddings, and illustrate the resulting visualizations with the Bernoulli (coin toss) problem, the ideal gas, $n$ sided die, the nonlinear least squares fit, and the Gaussian fit. We highlight how isKL can be used to determine the minimum number of parameters needed to describe probabilistic data, and conclude by visualizing the prediction space of the two-dimensional Ising model, where we examine the manifold behavior near its critical point.

cond-mat.stat-mech