arXiv · 2108.10724
How Hateful are Movies? A Study and Prediction on Movie Subtitles
Abstract
In this research, we investigate techniques to detect hate speech in movies. We introduce a new dataset collected from the subtitles of six movies, where each utterance is annotated either as hate, offensive or normal. We apply transfer learning techniques of domain adaptation and fine-tuning on existing social media datasets, namely from Twitter and Fox News. We evaluate different representations, i.e., Bag of Words (BoW), Bi-directional Long short-term memory (Bi-LSTM), and Bidirectional Encoder Representations from Transformers (BERT) on 11k movie subtitles. The BERT model obtained the best macro-averaged F1-score of 77%. Hence, we show that transfer learning from the social media domain is efficacious in classifying hate and offensive speech in movies through subtitles.
Explore related subjects
Keep this discovery
Niklas von Boguszewski, Sana Moin, Anirban Bhowmick, Seid Muhie Yimam, Chris Biemann. 2021-08-19. How Hateful are Movies? A Study and Prediction on Movie Subtitles. https://arxiv.org/abs/2108.10724
Cite the original work for its findings. Save a collection to share your selection of sources.