arXiv · 1508.07130
Parallel Dither and Dropout for Regularising Deep Neural Networks
Abstract
Effective regularisation during training can mean the difference between success and failure for deep neural networks. Recently, dither has been suggested as alternative to dropout for regularisation during batch-averaged stochastic gradient descent (SGD). In this article, we show that these methods fail without batch averaging and we introduce a new, parallel regularisation method that may be used without batch averaging. Our results for parallel-regularised non-batch-SGD are substantially better than what is possible with batch-SGD. Furthermore, our results demonstrate that dither and dropout are complimentary.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Andrew J. R. Simpson. 2015-08-28. Parallel Dither and Dropout for Regularising Deep Neural Networks. https://arxiv.org/abs/1508.07130
Cite the original work for its findings. Save a collection to share your selection of sources.