paper-with-me

홈 › Papers

Effective Black-Box Multi-Faceted Attacks Breach Vision Large Language Model Guardrails

2025-02-09 · Yijun Yang, Lichao Wang, Xiao Yang, Lanqing Hong, Jun Zhu

Vision Large Language Models (VLLMs) integrate visual data processing, expanding their real-world applications, but also increasing the risk of generating unsafe responses. In response, leading companies have implemented Multi-Layered safety defenses, including alignment training, safety system prompts, and content moderation. However, their effectiveness against sophisticated adversarial attacks remains largely unexplored. In this paper, we propose MultiFaceted Attack, a novel attack framework designed to systematically bypass Multi-Layered Defenses in VLLMs. It comprises three complementary attack facets: Visual Attack that exploits the multimodal nature of VLLMs to inject toxic system prompts through images; Alignment Breaking Attack that manipulates the model's alignment mechanism to prioritize the generation of contrasting responses; and Adversarial Signature that deceives content moderators by strategically placing misleading information at the end of the response. Extensive evaluations on eight commercial VLLMs in a black-box setting demonstrate that MultiFaceted Attack achieves a 61.56% attack success rate, surpassing state-of-the-art methods by at least 42.18%.

📄 PDF Abstract BibTeX arXiv:2502.05772

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

Acoustic Cybersecurity: Exploiting Voice-Activated Systems

2023-11-23 · Forrest McKee, David Noever

In this study, we investigate the emerging threat of inaudible acoustic attacks targeting digital voice assistants, a critical concern given their projected prevalence to exceed the global population by 2024. Our researc…

Text Embedding Inversion Security for Multilingual Language Models

2024-01-22 · Yiyi Chen, Heather Lent, Johannes Bjerva

Textual data is often represented as real-numbered embeddings in NLP, particularly with the popularity of large language models (LLMs) and Embeddings as a Service (EaaS). However, storing sensitive information as embeddi…

AutoBreach: Universal and Adaptive Jailbreaking with Efficient Wordplay-Guided Optimization

2024-05-30 · Jiawei Chen, Xiao Yang, Zhengwei Fang, Yu Tian 외

Despite the widespread application of large language models (LLMs) across various tasks, recent studies indicate that they are susceptible to jailbreak attacks, which can render their defense mechanisms ineffective. Howe…

SentenceSentence Compression

AutoAttacker: A Large Language Model Guided System to Implement Automatic Cyber-attacks

2024-03-02 · Jiacen Xu, Jack W. Stokes, Geoff McDonald, Xuesong Bai 외

Large language models (LLMs) have demonstrated impressive results on natural language tasks, and security researchers are beginning to employ them in both offensive and defensive systems. In cyber-security, there have be…

Computer SecurityLanguage ModelingLanguage ModellingLarge Language Model

Exploratory Analysis of Cyberattack Patterns on E-Commerce Platforms Using Statistical Methods

2025-11-04 · Fatimo Adenike Adeniya arxiv

Cyberattacks on e-commerce platforms have grown in sophistication, threatening consumer trust and operational continuity. This research presents a hybrid analytical framework that integrates statistical modelling and mac…

Ensemble Learning