paper-with-me

Papers

Sirens' Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMs

2026-03-14 · Zijian Ling, Pingyi Hu, Xiuyong Gao, Xiaojing Ma, Man Zhou, Jun Feng, Songfeng Lu, Dongmei Zhang, Bin Benjamin Zhu arxiv

Speech-driven large language models (LLMs) are increasingly accessed through speech interfaces, introducing new security risks via open acoustic channels. We present Sirens' Whisper (SWhisper), the first practical framework for covert prompt-based attacks against speech-driven LLMs under realistic black-box conditions using commodity hardware. SWhisper enables robust, inaudible delivery of arbitrary target baseband audio-including long and structured prompts-on commodity devices by encoding it into near-ultrasound waveforms that demodulate faithfully after acoustic transmission and microphone nonlinearity. This is achieved through a simple yet effective approach to modeling nonlinear channel characteristics across devices and environments, combined with lightweight channel-inversion pre-compensation. Building on this high-fidelity covert channel, we design a voice-aware jailbreak generation method that ensures intelligibility, brevity, and transferability under speech-driven interfaces. Experiments across both commercial and open-source speech-driven LLMs demonstrate strong black-box effectiveness. On commercial models, SWhisper achieves up to 0.94 non-refusal (NR) and 0.925 specific-convincing (SC). A controlled user study further shows that the injected jailbreak audio is perceptually indistinguishable from background-only playback for human listeners. Although jailbreaks serve as a case study, the underlying covert acoustic channel enables a broader class of high-fidelity prompt-injection and commandexecution attacks.

📄 PDF Abstract BibTeX arXiv:2603.13847

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

NUANCE: Near Ultrasound Attack On Networked Communication Environments

2023-04-25 · Forrest McKee, David Noever

This study investigates a primary inaudible attack vector on Amazon Alexa voice services using near ultrasound trojans and focuses on characterizing the attack surface and examining the practical implications of issuing …

Can You Hear It? Backdoor Attacks via Ultrasonic Triggers

2021-07-30 · Stefanos Koffas, Jing Xu, Mauro Conti, Stjepan Picek

This work explores backdoor attacks for automatic speech recognition systems where we inject inaudible triggers. By doing so, we make the backdoor attack challenging to detect for legitimate users, and thus, potentially …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Backdoor Attackspeech-recognition+1

Estimating Indoor Scene Depth Maps from Ultrasonic Echoes

2024-09-05 · Junpei Honma, Akisato Kimura, Go Irie

Measuring 3D geometric structures of indoor scenes requires dedicated depth sensors, which are not always available. Echo-based depth estimation has recently been studied as a promising alternative solution. All previous…

Depth Estimation

Adversarial Agents For Attacking Inaudible Voice Activated Devices

2023-07-23 · Forrest McKee, David Noever

The paper applies reinforcement learning to novel Internet of Thing configurations. Our analysis of inaudible attacks on voice-activated devices confirms the alarming risk factor of 7.6 out of 10, underlining significant…

CyberBattleSimQ-Learningreinforcement-learning

Robust Sensor Fusion Algorithms Against Voice Command Attacks in Autonomous Vehicles

2021-04-20 · Jiwei Guan, Xi Zheng, Chen Wang, Yipeng Zhou 외

With recent advances in autonomous driving, Voice Control Systems have become increasingly adopted as human-vehicle interaction methods. This technology enables drivers to use voice commands to control the vehicle and wi…

Autonomous DrivingAutonomous VehiclesMultimodal Deep LearningSensor Fusion