arXiv · 2404.05483
PetKaz at SemEval-2024 Task 8: Can Linguistics Capture the Specifics of LLM-generated Text?
Abstract
In this paper, we present our submission to the SemEval-2024 Task 8 "Multigenerator, Multidomain, and Multilingual Black-Box Machine-Generated Text Detection", focusing on the detection of machine-generated texts (MGTs) in English. Specifically, our approach relies on combining embeddings from the RoBERTa-base with diversity features and uses a resampled training set. We score 12th from 124 in the ranking for Subtask A (monolingual track), and our results show that our approach is generalizable across unseen models and domains, achieving an accuracy of 0.91.
Explore related subjects
Keep this discovery
Kseniia Petukhova, Roman Kazakov, Ekaterina Kochmar. 2024-04-08. PetKaz at SemEval-2024 Task 8: Can Linguistics Capture the Specifics of LLM-generated Text?. https://arxiv.org/abs/2404.05483
Cite the original work for its findings. Save a collection to share your selection of sources.