arXiv · 2602.18171
Click it or Leave it: Detecting and Spoiling Clickbait with Informativeness Measures and Large Language Models
Abstract
Clickbait headlines degrade the quality of online information and undermine user trust. We present a hybrid approach to clickbait detection that combines transformer-based text embeddings with linguistically motivated informativeness features. Using natural language processing techniques, we evaluate classical vectorizers, word embedding baselines, and large language model embeddings paired with tree-based classifiers. Our best-performing model, XGBoost over embeddings augmented with 15 explicit features, achieves an F1-score of 91\%, outperforming TF-IDF, Word2Vec, GloVe, LLM prompt based classification, and feature-only baselines. The proposed feature set enhances interpretability by highlighting salient linguistic cues such as second-person pronouns, superlatives, numerals, and attention-oriented punctuation, enabling transparent and well-calibrated clickbait predictions. We release code and trained models to support reproducible research.
Explore related subjects
Keep this discovery
Wojciech Michaluk, Tymoteusz Urban, Mateusz Kubita, Soveatin Kuntur, Anna Wroblewska. 2026-02-20. Click it or Leave it: Detecting and Spoiling Clickbait with Informativeness Measures and Large Language Models. https://arxiv.org/abs/2602.18171
Cite the original work for its findings. Save a collection to share your selection of sources.