KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
Jailbreak attacks exploit specific prompts to bypass LLM safeguards, causing the LLM to generate harmful, inappropriate, and misaligned content. Current jailbreaking methods rely heavily on carefully designed system prompts and numerous queries to achieve a single successful attack, which is costly and impractical for large-scale red-teaming. To address this challenge, we propose to distill the knowledge of an ensemble of SOTA attackers into a single open-source model, called Knowledge-Distilled Attacker (KDA), which is finetuned to automatically generate coherent and diverse attack prompts without the need for meticulous system prompt engineering. Compared to existing attackers, KDA achieves higher attack success rates and greater cost-time efficiency when targeting multiple SOTA open-source and commercial black-box LLMs. Furthermore, we conducted a quantitative diversity analysis of prompts generated by baseline methods and KDA, identifying diverse and ensemble attacks as key factors behind KDA's effectiveness and efficiency.
Code (0)
등록된 구현이 없습니다.
Tasks
DiversityPrompt EngineeringRed TeamingSimilar Papers 제목 키워드 기반
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
Automated red teaming can discover rare model failures and generate challenging examples that can be used for training or evaluation. However, a core challenge in automated red teaming is ensuring that the attacks are bo…
DiversityLarge Language ModelRed TeamingQuality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models
Ensuring safety of large language models (LLMs) is important. Red teaming--a systematic approach to identifying adversarial prompts that elicit harmful responses from target LLMs--has emerged as a crucial safety evaluati…
DiversityRed TeamingSentence EmbeddingSentence-EmbeddingMetaphor-based Jailbreak Attacks on Text-to-Image Models
Text-to-image (T2I) models commonly incorporate defense mechanisms to prevent the generation of sensitive images. Unfortunately, recent jailbreak attacks have shown that adversarial prompts can effectively bypass these m…
Disabling Self-Correction in Retrieval-Augmented Generation via Stealthy Retriever Poisoning
Retrieval-Augmented Generation (RAG) has become a standard approach for improving the reliability of large language models (LLMs). Prior work demonstrates the vulnerability of RAG systems by misleading them into generati…
HOI-PAGE: Zero-Shot Human-Object Interaction Generation with Part Affordance Guidance
We present HOI-PAGE, a new approach to synthesizing 4D human-object interactions (HOIs) from text prompts in a zero-shot fashion, driven by part-level affordance reasoning. In contrast to prior works that focus on global…
Human-Object Interaction DetectionHuman-Object Interaction GenerationObject