paper-with-me

Papers

JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models

2025-05-23 · Zifan Peng, Yule Liu, Zhen Sun, Mingchen Li, Zeren Luo, Jingyi Zheng, Wenhan Dong, Xinlei He, Xuechao Wang, Yingjie Xue, Shengmin Xu, Xinyi Huang

Audio Language Models (ALMs) have made significant progress recently. These models integrate the audio modality directly into the model, rather than converting speech into text and inputting text to Large Language Models (LLMs). While jailbreak attacks on LLMs have been extensively studied, the security of ALMs with audio modalities remains largely unexplored. Currently, there is a lack of an adversarial audio dataset and a unified framework specifically designed to evaluate and compare attacks and ALMs. In this paper, we present JALMBench, the \textit{first} comprehensive benchmark to assess the safety of ALMs against jailbreak attacks. JALMBench includes a dataset containing 2,200 text samples and 51,381 audio samples with over 268 hours. It supports 12 mainstream ALMs, 4 text-transferred and 4 audio-originated attack methods, and 5 defense methods. Using JALMBench, we provide an in-depth analysis of attack efficiency, topic sensitivity, voice diversity, and attack representations. Additionally, we explore mitigation strategies for the attacks at both the prompt level and the response level.

📄 PDF Abstract BibTeX arXiv:2505.17568

Code (1)

sfofgalaxy/jalmbench 공식 구현 pytorch

Tasks

BenchmarkingDiversity

Similar Papers 제목 키워드 기반

On Optimizing Multimodal Jailbreaks for Spoken Language Models

2026-03-19 · Aravind Krishnan, Karolina Stańczak, Dietrich Klakow arxiv

As Spoken Language Models (SLMs) integrate speech and text modalities, they inherit the safety vulnerabilities of their LLM backbone while introducing an expanded attack surface. SLMs have been previously shown to be sus…

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models

2025-05-21 · Zirui Song, Qian Jiang, Mingxuan Cui, Mingzhe Li 외

The rise of Large Audio Language Models (LAMs) brings both potential and risks, as their audio outputs may contain harmful or unethical content. However, current research lacks a systematic, quantitative evaluation of LA…

Bayesian OptimizationSpeech Synthesistext-to-speechText to Speech+1

Multilingual and Multi-Accent Jailbreaking of Audio LLMs

2025-04-01 · Jaechul Roh, Virat Shejwalkar, Amir Houmansadr

Large Audio Language Models (LALMs) have significantly advanced audio understanding but introduce critical security risks, particularly through audio jailbreaks. While prior work has focused on English-centric attacks, w…

GRM: Utility-Aware Jailbreak Attacks on Audio LLMs via Gradient-Ratio Masking

2026-04-10 · Yunqiang Wang, Hengyuan Na, Di Wu, Miao Hu 외 arxiv

Audio large language models (ALLMs) enable rich speech-text interaction, but they also introduce jailbreak vulnerabilities in the audio modality. Existing audio jailbreak methods mainly optimize jailbreak success while o…

Question Answering

Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs

2024-06-13 · Zhao Xu, Fan Liu, Hao liu

Although Large Language Models (LLMs) have demonstrated significant capabilities in executing complex tasks in a zero-shot manner, they are susceptible to jailbreak attacks and can be manipulated to produce harmful outpu…

BenchmarkingGPU