arXiv · cmp-lg/9506006
Automatic Extraction of Tagset Mappings from Parallel-Annotated Corpora
Abstract
This paper describes some of the recent work of project AMALGAM (automatic mapping among lexico-grammatical annotation models). We are investigating ways to map between the leading corpus annotation schemes in order to improve their resuability. Collation of all the included corpora into a single large annotated corpus will provide a more detailed language model to be developed for tasks such as speech and handwriting recognition. In particular, we focus here on a method of extracting mappings from corpora that have been annotated according to more than one annotation scheme.
Explore related subjects
Keep this discovery
John Hughes, Clive Souter, Eric Atwell. 1995-06-08. Automatic Extraction of Tagset Mappings from Parallel-Annotated Corpora. https://arxiv.org/abs/cmp-lg/9506006
Cite the original work for its findings. Save a collection to share your selection of sources.