arXiv · 2601.16890
LLM-Based Adversarial Persuasion Attacks on Fact-Checking Systems
Abstract
Automated fact-checking (AFC) systems are susceptible to adversarial attacks, enabling false claims to evade detection. Existing adversarial frameworks typically rely on injecting noise or altering semantics, yet no existing framework exploits the adversarial potential of persuasion techniques against AFC systems, which are widely used in disinformation campaigns to manipulate audiences. In this paper, we introduce a novel class of persuasive adversarial attacks on AFCs by employing an LLM to rephrase claims using persuasion techniques. Considering $15$ techniques grouped into $5$ categories, we study the effects of persuasion on both claim verification and evidence retrieval using a decoupled evaluation strategy. Experiments on the FEVER and FEVEROUS benchmarks show that persuasion attacks can substantially degrade both verification performance and evidence retrieval. Our analysis identifies persuasion techniques as a potent class of adversarial attacks, highlighting the need for more robust AFC systems.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
João A. Leite, Olesya Razuvayevskaya, Kalina Bontcheva, Carolina Scarton. 2026-01-23. LLM-Based Adversarial Persuasion Attacks on Fact-Checking Systems. https://arxiv.org/abs/2601.16890
Cite the original work for its findings. Save a collection to share your selection of sources.