paper-with-me

Papers

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs

2025-02-05 · Buyun Liang, Kwan Ho Ryan Chan, Darshan Thaker, Jinqi Luo, René Vidal

Jailbreak attacks exploit specific prompts to bypass LLM safeguards, causing the LLM to generate harmful, inappropriate, and misaligned content. Current jailbreaking methods rely heavily on carefully designed system prompts and numerous queries to achieve a single successful attack, which is costly and impractical for large-scale red-teaming. To address this challenge, we propose to distill the knowledge of an ensemble of SOTA attackers into a single open-source model, called Knowledge-Distilled Attacker (KDA), which is finetuned to automatically generate coherent and diverse attack prompts without the need for meticulous system prompt engineering. Compared to existing attackers, KDA achieves higher attack success rates and greater cost-time efficiency when targeting multiple SOTA open-source and commercial black-box LLMs. Furthermore, we conducted a quantitative diversity analysis of prompts generated by baseline methods and KDA, identifying diverse and ensemble attacks as key factors behind KDA's effectiveness and efficiency.

📄 PDF Abstract BibTeX arXiv:2502.05223

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityPrompt EngineeringRed Teaming

Similar Papers 제목 키워드 기반

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning

2024-12-24 · Alex Beutel, Kai Xiao, Johannes Heidecke, Lilian Weng

Automated red teaming can discover rare model failures and generate challenging examples that can be used for training or evaluation. However, a core challenge in automated red teaming is ensuring that the attacks are bo…

DiversityLarge Language ModelRed Teaming

Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models

2025-06-08 · Ren-Jian Wang, Ke Xue, Zeyu Qin, Ziniu Li 외

Ensuring safety of large language models (LLMs) is important. Red teaming--a systematic approach to identifying adversarial prompts that elicit harmful responses from target LLMs--has emerged as a crucial safety evaluati…

DiversityRed TeamingSentence EmbeddingSentence-Embedding

Metaphor-based Jailbreak Attacks on Text-to-Image Models

2025-12-06 · Chenyu Zhang, Lanjun Wang, Yiwen Ma, Wenhui Li 외 arxiv

Text-to-image (T2I) models commonly incorporate defense mechanisms to prevent the generation of sensitive images. Unfortunately, recent jailbreak attacks have shown that adversarial prompts can effectively bypass these m…

Disabling Self-Correction in Retrieval-Augmented Generation via Stealthy Retriever Poisoning

2025-08-27 · Yanbo Dai, Zhenlan Ji, Zongjie Li, Kuan Li 외 arxiv

Retrieval-Augmented Generation (RAG) has become a standard approach for improving the reliability of large language models (LLMs). Prior work demonstrates the vulnerability of RAG systems by misleading them into generati…

HOI-PAGE: Zero-Shot Human-Object Interaction Generation with Part Affordance Guidance

2025-06-08 · Lei LI, Angela Dai

We present HOI-PAGE, a new approach to synthesizing 4D human-object interactions (HOIs) from text prompts in a zero-shot fashion, driven by part-level affordance reasoning. In contrast to prior works that focus on global…

Human-Object Interaction DetectionHuman-Object Interaction GenerationObject