paper-with-me

Papers

Adversarial Attack-Defense Co-Evolution for LLM Safety Alignment via Tree-Group Dual-Aware Search and Optimization

2025-11-24 · Xurui Li, Kaisong Song, Rui Zhu, Pin-Yu Chen, Haixu Tang arxiv

Large Language Models (LLMs) have developed rapidly in web services, delivering unprecedented capabilities while amplifying societal risks. Existing works tend to focus on either isolated jailbreak attacks or static defenses, neglecting the dynamic interplay between evolving threats and safeguards in real-world web contexts. To mitigate these challenges, we propose ACE-Safety (Adversarial Co-Evolution for LLM Safety), a novel framework that jointly optimize attack and defense models by seamlessly integrating two key innovative procedures: (1) Group-aware Strategy-guided Monte Carlo Tree Search (GS-MCTS), which efficiently explores jailbreak strategies to uncover vulnerabilities and generate diverse adversarial samples; (2) Adversarial Curriculum Tree-aware Group Policy Optimization (AC-TGPO), which jointly trains attack and defense LLMs with challenging samples via curriculum reinforcement learning, enabling robust mutual improvement. Evaluations across multiple benchmarks demonstrate that our method outperforms existing attack and defense approaches, and provides a feasible pathway for developing LLMs that can sustainably support responsible AI ecosystems.

📄 PDF Abstract BibTeX arXiv:2511.19218

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningAdversarial Attack

Similar Papers 제목 키워드 기반

Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment

2026-05-03 · Jiajia Li, Xiaoyu Wen, Zhongtian Ma, Shuyue Hu 외 arxiv

The growing capabilities of large language models (LLMs) have driven their widespread deployment across diverse domains, even in potentially high-risk scenarios. Despite advances in safety alignment techniques, current m…

Co-Evolutionary Multi-Modal Alignment via Structured Adversarial Evolution

2026-03-02 · Guoxin Shi, Haoyu Wang, Zaihui Yang, Yuxing Wang 외 arxiv

Adversarial behavior plays a central role in aligning large language models with human values. However, existing alignment methods largely rely on static adversarial settings, which fundamentally limit robustness, partic…

MAGIC: A Co-Evolving Attacker-Defender Adversarial Game for Robust LLM Safety

2026-02-02 · Xiaoyu Wen, Zhida He, Han Qi, Ziyu Wan 외 arxiv

Ensuring robust safety alignment is crucial for Large Language Models (LLMs), yet existing defenses often lag behind evolving adversarial attacks due to their \textbf{reliance on static, pre-collected data distributions}…

Multi-agent Reinforcement Learning

Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution

2026-05-13 · Xiaozhe Zhang, Chaozhuo Li, Hui Liu, Shaocheng Yan 외 arxiv

Large language models remain vulnerable to adversarial prompts that elicit harmful outputs. Existing safety paradigms typically couple red-teaming and post-training in a closed, policy-centric loop, causing attack discov…

Red Teaming

Be Your Own Red Teamer: Safety Alignment via Self-Play and Reflective Experience Replay

2026-01-15 · Hao Wang, Yanting Wang, Hao Li, Rui Li 외 arxiv

Large Language Models (LLMs) have achieved remarkable capabilities but remain vulnerable to adversarial ``jailbreak'' attacks designed to bypass safety guardrails. Current safety alignment methods depend heavily on stati…

Reinforcement LearningRed Teaming