paper-with-me

홈 › Papers

JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs

2024-12-20 · Hongyi Li, Jiawei Ye, Jie Wu, Tianjie Yan, Chu Wang, Zhixin Li

Large Language Models (LLMs) aligned with human feedback have recently garnered significant attention. However, it remains vulnerable to jailbreak attacks, where adversaries manipulate prompts to induce harmful outputs. Exploring jailbreak attacks enables us to investigate the vulnerabilities of LLMs and further guides us in enhancing their security. Unfortunately, existing techniques mainly rely on handcrafted templates or generated-based optimization, posing challenges in scalability, efficiency and universality. To address these issues, we present JailPO, a novel black-box jailbreak framework to examine LLM alignment. For scalability and universality, JailPO meticulously trains attack models to automatically generate covert jailbreak prompts. Furthermore, we introduce a preference optimization-based attack method to enhance the jailbreak effectiveness, thereby improving efficiency. To analyze model vulnerabilities, we provide three flexible jailbreak patterns. Extensive experiments demonstrate that JailPO not only automates the attack process while maintaining effectiveness but also exhibits superior performance in efficiency, universality, and robustness against defenses compared to baselines. Additionally, our analysis of the three JailPO patterns reveals that attacks based on complex templates exhibit higher attack strength, whereas covert question transformations elicit riskier responses and are more likely to bypass defense mechanisms.

📄 PDF Abstract BibTeX arXiv:2412.15623

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BlackDAN: A Black-Box Multi-Objective Approach for Effective and Contextual Jailbreaking of Large Language Models

2024-10-13 · Xinyuan Wang, Victor Shea-Jay Huang, Renmiao Chen, Hao Wang 외

While large language models (LLMs) exhibit remarkable capabilities across various tasks, they encounter potential security risks such as jailbreak attacks, which exploit vulnerabilities to bypass security measures and ge…

Evolutionary Algorithms

Multi-Agent Framework for Threat Mitigation and Resilience in AI-Based Systems

2025-12-29 · Armstrong Foundjem, Lionel Nganyewou Tidjon, Leuson Da Silva, Foutse Khomh arxiv

Machine learning (ML) underpins foundation models in finance, healthcare, and critical infrastructure, making them targets for data poisoning, model extraction, prompt injection, automated jailbreaking, and preference-gu…

Model extraction

Robustness of Vision Language Models Against Split-Image Harmful Input Attacks

2026-02-08 · Md Rafi Ur Rashid, MD Sadik Hossain Shanto, Vishnu Asutosh Dasu, Shagufta Mehnaz arxiv

Vision-Language Models (VLMs) are now a core part of modern AI. Recent work proposed several visual jailbreak attacks using single/ holistic images. However, contemporary VLMs demonstrate strong robustness against such a…

Knowledge Distillation

Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search

2026-02-26 · Xun Huang, Simeng Qin, Xiaoshuang Jia, Ranjie Duan 외 arxiv

As Large Language Models (LLMs) are increasingly used, their security risks have drawn increasing attention. Existing research reveals that LLMs are highly susceptible to jailbreak attacks, with effectiveness varying acr…

GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMs

2024-11-21 · Advik Raj Basani, Xiao Zhang

LLMs have shown impressive capabilities across various natural language processing tasks, yet remain vulnerable to input prompts, known as jailbreak attacks, carefully designed to bypass safety guardrails and elicit harm…

Bayesian OptimizationRed Teaming