arXiv · 2103.12450
Are Neural Language Models Good Plagiarists? A Benchmark for Neural Paraphrase Detection
Abstract
The rise of language models such as BERT allows for high-quality text paraphrasing. This is a problem to academic integrity, as it is difficult to differentiate between original and machine-generated content. We propose a benchmark consisting of paraphrased articles using recent language models relying on the Transformer architecture. Our contribution fosters future research of paraphrase detection systems as it offers a large collection of aligned original and paraphrased documents, a study regarding its structure, classification experiments with state-of-the-art systems, and we make our findings publicly available.
Explore related subjects
Keep this discovery
Jan Philip Wahle, Terry Ruas, Norman Meuschke, Bela Gipp. 2021-03-23. Are Neural Language Models Good Plagiarists? A Benchmark for Neural Paraphrase Detection. https://doi.org/10.1109/jcdl52503.2021.00065
Cite the original work for its findings. Save a collection to share your selection of sources.