arXiv · 1707.06341
A Sub-Character Architecture for Korean Language Processing
Abstract
We introduce a novel sub-character architecture that exploits a unique compositional structure of the Korean language. Our method decomposes each character into a small set of primitive phonetic units called jamo letters from which character- and word-level representations are induced. The jamo letters divulge syntactic and semantic information that is difficult to access with conventional character-level units. They greatly alleviate the data sparsity problem, reducing the observation space to 1.6% of the original while increasing accuracy in our experiments. We apply our architecture to dependency parsing and achieve dramatic improvement over strong lexical baselines.
Explore related subjects
Keep this discovery
Karl Stratos. 2017-07-20. A Sub-Character Architecture for Korean Language Processing. https://arxiv.org/abs/1707.06341
Cite the original work for its findings. Save a collection to share your selection of sources.