arXiv · 2301.11487
Projected Subnetworks Scale Adaptation
Abstract
Large models support great zero-shot and few-shot capabilities. However, updating these models on new tasks can break performance on previous seen tasks and their zero/few-shot unseen tasks. Our work explores how to update zero/few-shot learners such that they can maintain performance on seen/unseen tasks of previous tasks as well as new tasks. By manipulating the parameter updates of a gradient-based meta learner as the projected task-specific subnetworks, we show improvements for large models to retain seen and zero/few shot task performance in online settings.
Explore related subjects
Keep this discovery
Siddhartha Datta, Nigel Shadbolt. 2023-01-27. Projected Subnetworks Scale Adaptation. https://arxiv.org/abs/2301.11487
Cite the original work for its findings. Save a collection to share your selection of sources.