arXiv · 2410.07041
Emergent properties with repeated examples
Abstract
We study the performance of transformers as a function of the number of repetitions of training examples with algorithmically generated datasets. On three problems of mathematics: the greatest common divisor, modular multiplication, and matrix eigenvalues, we show that for a fixed number of training steps, models trained on smaller sets of repeated examples outperform models trained on larger sets of single-use examples. We also demonstrate that two-set training - repeated use of a small random subset of examples, along normal sampling on the rest of the training set - provides for faster learning and better performance. This highlights that the benefits of repetition can outweigh those of data diversity. These datasets and problems provide a controlled setting to shed light on the still poorly understood interplay between generalization and memorization in deep learning.
Explore related subjects
Keep this discovery
François Charton, Julia Kempe. 2024-10-09. Emergent properties with repeated examples. https://arxiv.org/abs/2410.07041
Cite the original work for its findings. Save a collection to share your selection of sources.