paper-with-me

Papers

LLM-Based Adversarial Persuasion Attacks on Fact-Checking Systems

2026-01-23 · João A. Leite, Olesya Razuvayevskaya, Kalina Bontcheva, Carolina Scarton arxiv

Automated fact-checking (AFC) systems are susceptible to adversarial attacks, enabling false claims to evade detection. Existing adversarial frameworks typically rely on injecting noise or altering semantics, yet no existing framework exploits the adversarial potential of persuasion techniques, which are widely used in disinformation campaigns to manipulate audiences. In this paper, we introduce a novel class of persuasive adversarial attacks on AFCs by employing a generative LLM to rephrase claims using persuasion techniques. Considering 15 techniques grouped into 6 categories, we study the effects of persuasion on both claim verification and evidence retrieval using a decoupled evaluation strategy. Experiments on the FEVER and FEVEROUS benchmarks show that persuasion attacks can substantially degrade both verification performance and evidence retrieval. Our analysis identifies persuasion techniques as a potent class of adversarial attacks, highlighting the need for more robust AFC systems.

📄 PDF Abstract BibTeX arXiv:2601.16890

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring

2026-07-09 · Jennifer Za, Julija Bainiaksina, Nikita Ostrovsky, Tanush Chopra 외 arxiv

Chain-of-thought (CoT) monitoring is a promising safety mechanism for AI agents, based on the premise that visible reasoning traces can surface misaligned or deceptive behavior. While effective in standard scenarios, rec…

DECEIVE-AFC: Adversarial Claim Attacks against Search-Enabled LLM-based Fact-Checking Systems

2026-01-31 · Haoran Ou, Kangjie Chen, Gelei Deng, Hangcheng Liu 외 arxiv

Fact-checking systems with search-enabled large language models (LLMs) have shown strong potential for verifying claims by dynamically retrieving external evidence. However, the robustness of such systems against adversa…

Adversarial Attack

Adversarial Attacks Against Automated Fact-Checking: A Survey

2025-09-10 · Fanzhen Liu, Alsharif Abuadbba, Kristen Moore, Surya Nepal 외 arxiv

In an era where misinformation spreads freely, fact-checking (FC) plays a crucial role in verifying claims and promoting reliable information. While automated fact-checking (AFC) has advanced significantly, existing syst…

Generating Label Cohesive and Well-Formed Adversarial Claims

2020-09-17 · EMNLP 2020 11 · Pepa Atanasova, Dustin Wright, Isabelle Augenstein

Adversarial attacks reveal important vulnerabilities and flaws of trained models. One potent type of attack are universal adversarial triggers, which are individual n-grams that, when appended to instances of a class und…

Fact CheckingLanguage ModelingLanguage ModellingNatural Language Inference+1

Synthetic Disinformation Attacks on Automated Fact Verification Systems

2022-02-18 · Yibing Du, Antoine Bosselut, Christopher D. Manning

Automated fact-checking is a needed technology to curtail the spread of online misinformation. One current framework for such solutions proposes to verify claims by retrieving supporting or refuting evidence from related…

Fact CheckingFact VerificationMisinformation