paper-with-me

Papers

RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue

2026-09-15 · Yuqi Wang, Fengyuan Liu, Haochen Luo, Zhiqi Yu, Qi Liu arxiv

Speech-to-speech dialogue models increasingly support persona control, yet existing spoken role-playing benchmarks remain largely character-centric and short-horizon. This leaves open whether spoken dialogue models can sustain diverse roles over extended interactions, especially beyond predefined fictional characters. We introduce RoleBreak, an open benchmark for long-horizon role-playing robustness in spoken dialogue. RoleBreak contains 310 character-based and user-centered roles, 6,688 human-verified dialogue turns, and 11,743 fine-grained evaluation criteria, with 1,856 turns carrying expressive emotion targets for evaluating vocal emotion. Its scenarios are designed to stress role consistency, interaction quality, safety, and affect over extended conversations. We evaluate nine configurations spanning full-duplex, omni-modal, and cascaded ASR--LLM--TTS paradigms. We find four key patterns. First, current systems are substantially stronger at semantic role adherence than at vocal emotion. Second, semantic robustness remains brittle over long interactions: even the strongest evaluated system encounters its first persona and safety failures after only 10.4 and 11.6 turns on average. Third, scaling the LLM substantially improves semantic robustness and delays failure, but yields little improvement in vocal emotion. Finally, user vocal emotion affects role-playing behavior even when linguistic content is fixed. These findings highlight persistent gaps in both long-horizon robustness and vocal expressiveness in spoken role-playing systems.

📄 PDF Abstract BibTeX arXiv:2609.16614

Code (3)

InsomaniacElf/sg-tamil-tts-resources- ★ 1
Tavish9/awesome-daily-AI-arxiv ★ 117
iszhanjiawei/TTS_arxiv_daily ★ 3

Similar Papers 제목 키워드 기반

RoleBreak: Character Hallucination as a Jailbreak Attack in Role-Playing Systems

2024-09-25 · Yihong Tang, Bo wang, Xu Wang, Dongming Zhao 외

Role-playing systems powered by large language models (LLMs) have become increasingly influential in emotional communication applications. However, these systems are susceptible to character hallucinations, where the mod…

Hallucination

RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models

2023-10-01 · Zekun Moore Wang, Zhongyuan Peng, Haoran Que, Jiaheng Liu 외

The advent of Large Language Models (LLMs) has paved the way for complex tasks such as role-playing, which enhances user interactions by enabling models to imitate various characters. However, the closed-source nature of…

Benchmarking

DynSess: Dynamic Session-Level Evaluation and Optimization Framework for Role-Playing Agents

2026-05-28 · Rongsheng Zhang, Jiji Tang, Junnan Ren, Zuyi Bao 외 arxiv

Role-playing with large language models is fundamentally a session-level task, requiring agents to sustain character identity and interaction quality across extended multi-turn conversations. Yet existing evaluation and …

Benchmarking Bias in Large Language Models during Role-Playing

2024-11-01 · Xinyue Li, Zhenpeng Chen, Jie M. Zhang, Yiling Lou 외

Large Language Models (LLMs) have become foundational in modern language-driven applications, profoundly influencing daily life. A critical technique in leveraging their potential is role-playing, where LLMs simulate div…

BenchmarkingFairnessMultiple-choice

SAGED: A Holistic Bias-Benchmarking Pipeline for Language Models with Customisable Fairness Calibration

2024-09-17 · Xin Guan, Ze Wang, Nathaniel Demchak, Saloni Gupta 외

The development of unbiased large language models is widely recognized as crucial, yet existing benchmarks fall short in detecting biases due to limited scope, contamination, and lack of a fairness baseline. SAGED(bias) …

BenchmarkingcounterfactualFairnessSentiment Analysis