paper-with-me

Papers

StyleBench: Evaluating Speech Language Models on Conversational Speaking Style Control

2026-03-08 · Haishu Zhao, Aokai Hao, Yuan Ge, Zhenqiang Hong, Tong Xiao, Jingbo Zhu arxiv

Speech language models (SLMs) have significantly extended the interactive capability of text-based Large Language Models (LLMs) by incorporating paralinguistic information. For more realistic interactive experience with customized styles, current SLMs have managed to interpret and control speaking style intensity from user prompts during the dialogue process. However, there remains a lack of systematic benchmarks that quantifies and evaluates the style intensity control ability in conversations. In this paper, we propose StyleBench, a multi-turn dialogue benchmark for comprehensively evaluating the style intensity control ability across four dimensions: emotion, speed, volume, and pitch. Our results reveal the performance gaps between leading SLMs and omni language models (OLMs), suggesting the underlying reasons and promising approaches for future exploration.

📄 PDF Abstract BibTeX arXiv:2603.07599

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Enhancing Speaking Styles in Conversational Text-to-Speech Synthesis with Graph-based Multi-modal Context Modeling

2021-06-11 · Jingbei Li, Yi Meng, Chenyi Li, Zhiyong Wu 외

Comparing with traditional text-to-speech (TTS) systems, conversational TTS systems are required to synthesize speeches with proper speaking style confirming to the conversational context. However, state-of-the-art conte…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Development of Hybrid ASR Systems for Low Resource Medical Domain Conversational Telephone Speech

2022-10-24 · Christoph Lüscher, Mohammad Zeineldeen, Zijian Yang, Tina Raissi 외

Language barriers present a great challenge in our increasingly connected and global world. Especially within the medical domain, e.g. hospital or emergency room, communication difficulties and delays may lead to malprac…

automatic-speech-translationTranslation

Audiobook Dialogues as Training Data for Conversational Style Synthetic Voices

2022-06-01 · LREC 2022 6 · Liisi Piits, Hille Pajupuu, Heete Sahkai, Rene Altrov 외

Synthetic voices are increasingly used in applications that require a conversational speaking style, raising the question as to which type of training data yields the most suitable speaking style for such applications. T…

Sentencetext-to-speechText to Speech

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models

2025-09-30 · Kai-Wei Chang, En-Pei Hu, Chun-Yi Kuan, Wenze Ren 외 arxiv

Conversational Spoken Language Models (SLMs) are emerging as a promising paradigm for real-time speech interaction. However, their capacity of temporal dynamics, including the ability to manage timing, tempo and simultan…

Turn-Taking Prediction for Natural Conversational Speech

2022-08-29 · Shuo-Yiin Chang, Bo Li, Tara N. Sainath, Chao Zhang 외

While a streaming voice assistant system has been used in many applications, this system typically focuses on unnatural, one-shot interactions assuming input from a single voice query without hesitation or disfluency. Ho…

Predictionspeech-recognitionSpeech Recognition