paper-with-me

Papers

Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models

2024-05-09 · Vyas Raina, Rao Ma, Charles McGhee, Kate Knill, Mark Gales

Recent developments in large speech foundation models like Whisper have led to their widespread use in many automatic speech recognition (ASR) applications. These systems incorporate special tokens' in their vocabulary, such as $\texttt{<|endoftext|>}$, to guide their language generation process. However, we demonstrate that these tokens can be exploited by adversarial attacks to manipulate the model's behavior. We propose a simple yet effective method to learn a universal acoustic realization of Whisper's $\texttt{<|endoftext|>}$ token, which, when prepended to any speech signal, encourages the model to ignore the speech and only transcribe the special token, effectively muting' the model. Our experiments demonstrate that the same, universal 0.64-second adversarial audio segment can successfully mute a target Whisper ASR model for over 97\% of speech samples. Moreover, we find that this universal adversarial audio segment often transfers to new datasets and tasks. Overall this work demonstrates the vulnerability of Whisper models to `muting' adversarial attacks, where such attacks can pose both risks and potential benefits in real-world settings: for example the attack can be used to bypass speech moderation systems, or conversely the attack can also be used to protect private speech data.

📄 PDF Abstract BibTeX arXiv:2405.06134

Code (1)

rainavyas/prepend_acoustic_attack 공식 구현 pytorch

Tasks

Adversarial AttackAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionText Generation

Similar Papers 제목 키워드 기반

Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models

2024-07-05 · Vyas Raina, Mark Gales

Speech enabled foundation models, either in the form of flexible speech recognition based systems or audio-prompted large language models (LLMs), are becoming increasingly popular. One of the interesting aspects of these…

Adversarial AttackAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Form+3

Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs

2025-05-20 · Rao Ma, Mengjie Qian, Vyas Raina, Mark Gales 외

The combination of pre-trained speech encoders with large language models has enabled the development of speech LLMs that can handle a wide range of spoken language processing tasks. While these models are powerful and f…

Attribute

Room for Error: Large-Scale Simulation of Over-the-Air Acoustic Attacks

2026-06-26 · Andrew C. Cullen, Neil G. Marchant, Jiani Xie, Paul Montague 외 arxiv

While voice control is rapidly becoming a ubiquitous vector of human-AI communication, the risks facing these systems remain poorly understood. This is, in part, a product of the difficulties in scaling strictly digital …

Sirens' Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMs

2026-03-14 · Zijian Ling, Pingyi Hu, Xiuyong Gao, Xiaojing Ma 외 arxiv

Speech-driven large language models (LLMs) are increasingly accessed through speech interfaces, introducing new security risks via open acoustic channels. We present Sirens' Whisper (SWhisper), the first practical framew…

Multilingual and Multi-Accent Jailbreaking of Audio LLMs

2025-04-01 · Jaechul Roh, Virat Shejwalkar, Amir Houmansadr

Large Audio Language Models (LALMs) have significantly advanced audio understanding but introduce critical security risks, particularly through audio jailbreaks. While prior work has focused on English-centric attacks, w…