arXiv · 2404.17651
Hard ASH: Sparsity and the right optimizer make a continual learner
Abstract
In class incremental learning, neural networks typically suffer from catastrophic forgetting. We show that an MLP featuring a sparse activation function and an adaptive learning rate optimizer can compete with established regularization techniques in the Split-MNIST task. We highlight the effectiveness of the Adaptive SwisH (ASH) activation function in this context and introduce a novel variant, Hard Adaptive SwisH (Hard ASH) to further enhance the learning retention.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Santtu Keskinen. 2024-04-26. Hard ASH: Sparsity and the right optimizer make a continual learner. https://arxiv.org/abs/2404.17651
Cite the original work for its findings. Save a collection to share your selection of sources.