arXiv · 2105.06514
Distilling BERT for low complexity network training
Abstract
This paper studies the efficiency of transferring BERT learnings to low complexity models like BiLSTM, BiLSTM with attention and shallow CNNs using sentiment analysis on SST-2 dataset. It also compares the complexity of inference of the BERT model with these lower complexity models and underlines the importance of these techniques in enabling high performance NLP models on edge devices like mobiles, tablets and MCU development boards like Raspberry Pi etc. and enabling exciting new applications.
Explore related subjects
Keep this discovery
Bansidhar Mangalwedhekar. 2021-05-13. Distilling BERT for low complexity network training. https://arxiv.org/abs/2105.06514
Cite the original work for its findings. Save a collection to share your selection of sources.