arXiv · 1912.00544
Multi-Scale Self-Attention for Text Classification
Abstract
In this paper, we introduce the prior knowledge, multi-scale structure, into self-attention modules. We propose a Multi-Scale Transformer which uses multi-scale multi-head self-attention to capture features from different scales. Based on the linguistic perspective and the analysis of pre-trained Transformer (BERT) on a huge corpus, we further design a strategy to control the scale distribution for each layer. Results of three different kinds of tasks (21 datasets) show our Multi-Scale Transformer outperforms the standard Transformer consistently and significantly on small and moderate size datasets.
Explore related subjects
Keep this discovery
Qipeng Guo, Xipeng Qiu, Pengfei Liu, Xiangyang Xue, Zheng Zhang. 2019-12-02. Multi-Scale Self-Attention for Text Classification. https://arxiv.org/abs/1912.00544
Cite the original work for its findings. Save a collection to share your selection of sources.