arXiv · 1910.02586
Why Attention? Analyzing and Remedying BiLSTM Deficiency in Modeling Cross-Context for NER
Abstract
State-of-the-art approaches of NER have used sequence-labeling BiLSTM as a core module. This paper formally shows the limitation of BiLSTM in modeling cross-context patterns. Two types of simple cross-structures -- self-attention and Cross-BiLSTM -- are shown to effectively remedy the problem. On both OntoNotes 5.0 and WNUT 2017, clear and consistent improvements are achieved over bare-bone models, up to 8.7% on some of the multi-token mentions. In-depth analyses across several aspects of the improvements, especially the identification of multi-token mentions, are further given.
Explore related subjects
Keep this discovery
Peng-Hsuan Li, Tsu-Jui Fu, Wei-Yun Ma. 2019-10-07. Why Attention? Analyzing and Remedying BiLSTM Deficiency in Modeling Cross-Context for NER. https://arxiv.org/abs/1910.02586
Cite the original work for its findings. Save a collection to share your selection of sources.