arXiv · 1909.12131
MinWikiSplit: A Sentence Splitting Corpus with Minimal Propositions
Abstract
We compiled a new sentence splitting corpus that is composed of 203K pairs of aligned complex source and simplified target sentences. Contrary to previously proposed text simplification corpora, which contain only a small number of split examples, we present a dataset where each input sentence is broken down into a set of minimal propositions, i.e. a sequence of sound, self-contained utterances with each of them presenting a minimal semantic unit that cannot be further decomposed into meaningful propositions. This corpus is useful for developing sentence splitting approaches that learn how to transform sentences with a complex linguistic structure into a fine-grained representation of short sentences that present a simple and more regular structure which is easier to process for downstream applications and thus facilitates and improves their performance.
Explore related subjects
Keep this discovery
Christina Niklaus, Andre Freitas, Siegfried Handschuh. 2019-09-26. MinWikiSplit: A Sentence Splitting Corpus with Minimal Propositions. https://arxiv.org/abs/1909.12131
Cite the original work for its findings. Save a collection to share your selection of sources.