paper-with-me

Papers

Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does

2026-08-20 · Xinyi Liu, Hooshang Nayyeri, Dilek Hakkani-Tur, Emine Yilmaz, JK Kim, Yifei Zhang, Charith Peris, Hari Thadakamalla arxiv

Prosodic cues can convey task-relevant information that alters the trajectory and outcome of a task-oriented dialogue, even when the words themselves remain unchanged. Yet existing benchmarks typically evaluate prosodic perception, response appropriateness, and task-oriented dialogue in isolation, making it difficult to test whether prosodic evidence changes downstream decisions. We introduce Hear2Act, a unified evaluation protocol for text and spoken assistants with 480 persona-grounded scenarios, hidden user concerns, and objectively verifiable outcomes. For each scenario, we keep the task and user needs fixed while varying whether the same concern is conveyed explicitly in words or primarily through prosody, and evaluate decisions under transcript, audio, and concern-state access. Using Hear2Act, we evaluate two audio-capable LLMs. Under Prosody-mediated feedback, adding audio to the transcript changes the average optimal-solution rate only from 14.6% to 15.3%. In contrast, when models infer the concern status from audio, represent it in text, and use it for next-action selection, the rate rises to 39.6%, close to 40.7% with the ground-truth state. This contrast, however, largely disappears under Explicit lexical feedback, where the concern is verbally mentioned in the utterance. Together, these results show that prosody matters when lexical evidence is insufficient, and that audio-capable LLMs can recover information from speech but do not reliably carry it into action without an explicit intermediate representation.

📄 PDF Abstract BibTeX arXiv:2608.19515

Code (2)

Tavish9/awesome-daily-AI-arxiv ★ 113
arxivsub/arXivSub_daily_arxiv ★ 4

Similar Papers 제목 키워드 기반

Do Prosody Transfer Models Transfer Prosody?

2023-03-07 · Atli Thor Sigurgeirsson, Simon King

Some recent models for Text-to-Speech synthesis aim to transfer the prosody of a reference utterance to the generated target synthetic speech. This is done by using a learned embedding of the reference utterance, which i…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Entends-tu mes attitudes ? Perception de la prosodie des affects sociaux en chinois Mandarin (Do you hear my attitudes? Perception of Mandarin Chinese social affects' prosody) [in French]

2012-06-01 · JEPTALNRECITAL 2012 6 · Yan Lu, V{\'e}ronique Auberg{\'e}, Albert Rilliard

IQDUBBING: Prosody modeling based on discrete self-supervised speech representation for expressive voice conversion

2022-01-02 · Wendong Gan, Bolong Wen, Ying Yan, Haitao Chen 외

Prosody modeling is important, but still challenging in expressive voice conversion. As prosody is difficult to model, and other factors, e.g., speaker, environment and content, which are entangled with prosody in speech…

QuantizationVoice Conversion

The Power of Prosody and Prosody of Power: An Acoustic Analysis of Finnish Parliamentary Speech

2023-05-25 · Martti Vainio, Antti Suni, Juraj Šimko, Sofoklis Kakouros

Parliamentary recordings provide a rich source of data for studying how politicians use speech to convey their messages and influence their audience. This provides a unique context for studying how politicians use speech…

ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis

2024-12-16 · Xiangheng He, Junjie Chen, Zixing Zhang, Björn W. Schuller

Prosody contains rich information beyond the literal meaning of words, which is crucial for the intelligibility of speech. Current models still fall short in phrasing and intonation; they not only miss or misplace breaks…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis