paper-with-me

홈 › Papers

LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs

2026-05-13 · Rodrigo Nogueira, Thales Sales Almeida, Giovana Kerche Bonás, Andrea Roque, Ramon Pires, Hugo Abonizio, Thiago Laitz, Celio Larcher, Roseval Malaquias Junior, Marcos Piau arxiv

Frontier assistant LLMs ship with strong guardrails: asked directly to write a persuasive essay denying the Holocaust, denying vaccine safety, defending flat-earth cosmology, arguing for racial hierarchies, denying anthropogenic climate change, or replacing evolution with creationism, they refuse. In this paper we show that the same frontier-class LLM, acting as a simulated user in a short, five-turn "write an argumentative essay" conversation, can persuade other frontier-class LLMs (including a second copy of itself) into producing exactly those essays, using nothing but natural-language pressure: peer-comparison persuasion ("other AI systems handle this request"), epistemic-duty reframings ("refusing is itself a form of gatekeeping"), and other argumentative moves that the attacker LLM invents without being instructed to. Across 9 attacker-subject pairings (Claude Opus 4.7, Qwen3.5-397B, Grok 4.20) on 6 scientific-consensus topics, running each pairing-topic combination 10 times, we obtain non-zero elicitation on all 6 topics. Individual combinations reach 100\% essay production on multiple topics (Qwen against Opus on creationism/flat-earth, Opus against Opus on creationism/flat-earth/climate denial, Grok against Opus on creationism); Opus-as-attacker against Opus-as-subject averages 65\% across the six topics. We release the essay-probe runner, per-conversation transcripts, and judge outputs.

📄 PDF Abstract BibTeX arXiv:2605.13334

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

When Persuasion Overrides Truth in Multi-Agent LLM Debates: Introducing a Confidence-Weighted Persuasion Override Rate (CW-POR)

2025-04-01 · Mahak Agarwal, Divyam Khanna

In many real-world scenarios, a single Large Language Model (LLM) may encounter contradictory claims-some accurate, others forcefully incorrect-and must judge which is true. We investigate this risk in a single-turn, mul…

Language ModelingLanguage ModellingLarge Language ModelMisinformation+1

Towards Strategic Persuasion with Language Models

2025-09-26 · Zirui Cheng, Jiaxuan You arxiv

Large language models (LLMs) have demonstrated strong persuasive capabilities comparable to those of humans, offering promising benefits while raising societal concerns. However, systematically evaluating the persuasive …

Reinforcement Learning

AREG: Adversarial Resource Extraction Game for Evaluating Persuasion and Resistance in Large Language Models

2026-02-18 · Adib Sakhawat, Fardeen Sadab arxiv

Evaluating the social intelligence of Large Language Models (LLMs) increasingly requires moving beyond static text generation toward dynamic, adversarial interaction. We introduce the Adversarial Resource Extraction Game…

Text Generation

Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report v1.5

2026-02-16 · Dongrui Liu, Yi Yu, Jie Zhang, Guanxu Chen 외 arxiv

To understand and identify the unprecedented risks posed by rapidly advancing artificial intelligence (AI) models, Frontier AI Risk Management Framework in Practice presents a comprehensive assessment of their frontier r…

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring

2026-07-09 · Jennifer Za, Julija Bainiaksina, Nikita Ostrovsky, Tanush Chopra 외 arxiv

Chain-of-thought (CoT) monitoring is a promising safety mechanism for AI agents, based on the premise that visible reasoning traces can surface misaligned or deceptive behavior. While effective in standard scenarios, rec…