paper-with-me

Papers

Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis

2026-07-29 · Jiachen Qian, Junyu Li arxiv

Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains unclear how much jailbreak capability can arise from matched-text variation in speech delivery rather than from lexical rewriting or broader style transfer. We study this question by holding transcript content fixed and varying six speech-delivery presets whose acoustic attributes may co-vary. We present PJ-Break, a black-box evaluation protocol with presets targeting arousal, authority, and speaking rate, together with AdvAudio-Prosody, a 600-sample benchmark with acoustically verified attributes. On the exact post-QC Qwen2-Audio panel, the Q=1 Panic (38/95), Anger (35/95), and Fast (32/95) presets are all well above Neutral (4/95). The fixed six-query pool covers 44/95 Qwen2-Audio seeds and 15/95 GPT-4o seeds and exceeds a matched-budget StyleBreak reimplementation (27/95) on Qwen2-Audio. A same-voice pool excluding the confounded Commanding condition still reaches 40/95, and a retained-panel ablation shows emotional-delivery audio alone (44/95) is far more effective than emotional text alone (11/95). Exploratory surrogate diagnostics and pilot mitigation observations are secondary, non-core analyses. Overall, matched-text speech delivery should be treated as a first-class factor in Audio LLM safety evaluation

📄 PDF Abstract BibTeX arXiv:2607.26541

Code (0)

등록된 구현이 없습니다.

Tasks

Style Transfer

Similar Papers 제목 키워드 기반

Sirens' Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMs

2026-03-14 · Zijian Ling, Pingyi Hu, Xiuyong Gao, Xiaojing Ma 외 arxiv

Speech-driven large language models (LLMs) are increasingly accessed through speech interfaces, introducing new security risks via open acoustic channels. We present Sirens' Whisper (SWhisper), the first practical framew…

DEAF: A Benchmark for Diagnostic Evaluation of Acoustic Faithfulness in Audio Language Models

2026-03-17 · Jiaqi Xiong, Yunjia Qi, Qi Cao, Yu Zheng 외 arxiv

Recent Audio Multimodal Large Language Models (Audio MLLMs) demonstrate impressive performance on speech benchmarks, yet it remains unclear whether these models genuinely process acoustic signals or rely on text-based se…

Hierarchical Prosody Modeling for Non-Autoregressive Speech Synthesis

2020-11-12 · Chung-Ming Chien, Hung-Yi Lee

Prosody modeling is an essential component in modern text-to-speech (TTS) frameworks. By explicitly providing prosody features to the TTS model, the style of synthesized utterances can thus be controlled. However, predic…

Speech Synthesistext-to-speechText to Speech

Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does

2026-08-20 · Xinyi Liu, Hooshang Nayyeri, Dilek Hakkani-Tur, Emine Yilmaz 외 arxiv

Prosodic cues can convey task-relevant information that alters the trajectory and outcome of a task-oriented dialogue, even when the words themselves remain unchanged. Yet existing benchmarks typically evaluate prosodic …

Improving Prosody Modelling with Cross-Utterance BERT Embeddings for End-to-end Speech Synthesis

2020-11-06 · Guanghui Xu, Wei Song, Zhengchen Zhang, Chao Zhang 외

Despite prosody is related to the linguistic information up to the discourse structure, most text-to-speech (TTS) systems only take into account that within each sentence, which makes it challenging when converting a par…

DecoderSentenceSentence EmbeddingsSpeech Synthesis+2