paper-with-me

홈 › Papers

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents

2025-09-04 · Weihao Wu, Liang Cao, Xinyu Wu, Zhiwei Lin, Rui Niu, Jingbei Li, Zhiyong Wu arxiv

Recent significant advancements in Large Language Models (LLMs) have greatly propelled the development of Role-Playing Conversational Agents (RPCAs). These systems aim to create immersive user experiences through consistent persona adoption. However, current RPCA research faces dual limitations. First, existing work predominantly focuses on the textual modality, entirely overlooking critical paralinguistic features including intonation, prosody, and rhythm in speech, which are essential for conveying character emotions and shaping vivid identities. Second, the speech-based role-playing domain suffers from a long-standing lack of standardized evaluation benchmarks. Most current spoken dialogue datasets target only fundamental capability assessments, featuring thinly sketched or ill-defined character profiles. Consequently, they fail to effectively quantify model performance on core competencies like long-term persona consistency. To address this critical gap, we introduce VoxRole, the first comprehensive benchmark specifically designed for the evaluation of speech-based RPCAs. The benchmark comprises 13335 multi-turn dialogues, totaling 65.6 hours of speech from 1228 unique characters across 261 movies. To construct this resource, we propose a novel two-stage automated pipeline that first aligns movie audio with scripts and subsequently employs an LLM to systematically build multi-dimensional profiles for each character. Leveraging VoxRole, we conduct a multi-dimensional evaluation of contemporary spoken dialogue models, revealing crucial insights into their respective strengths and limitations in maintaining persona consistency.

📄 PDF Abstract BibTeX arXiv:2509.03940

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SpeechRole: A Large-Scale Dataset and Benchmark for Evaluating Speech Role-Playing Agents

2025-08-04 · Changhao Jiang, Jiajun Sun, Yifei Cao, Jiabao Zhuang 외 arxiv

Speech is essential for realistic role-playing, yet existing work on role-playing agents largely centers on text, leaving Speech Role-Playing Agents (SRPAs) underexplored and without systematic evaluation. We introduce S…

KoALa-Bench: Evaluating Large Audio Language Models on Korean Speech Understanding and Faithfulness

2026-03-30 · Jinyoung Kim, Hyeongsoo Lim, Eunseo Seo, Minho Jang 외 arxiv

Recent advances in large audio language models (LALMs) have enabled multilingual speech understanding. However, benchmarks for evaluating LALMs remain scarce for non-English languages, with Korean being one such underexp…

Instruction FollowingSpeech RecognitionQuestion Answering

Speech-DRAME: A Framework for Human-Aligned Benchmarks in Speech Role-Play

2025-11-03 · Jiatong Shi, Jionghao Han, Yichen Lu, Santiago Pascual 외 arxiv

Role-play has become a key testbed for generative models, expanding from text-only dialogue to multimodal interaction. Extending role-play to speech captures prosody, emotion, and delivery, but also poses new evaluation …

Benchmarking Representations for Speech, Music, and Acoustic Events

2024-05-02 · Moreno La Quatra, Alkis Koudounas, Lorenzo Vaiani, Elena Baralis 외

Limited diversity in standardized benchmarks for evaluating audio representation learning (ARL) methods may hinder systematic comparison of current methods' capabilities. We present ARCH, a comprehensive benchmark for ev…

Audio ClassificationBenchmarkingDiversityRepresentation Learning

To Dereverb Or Not to Dereverb? Perceptual Studies On Real-Time Dereverberation Targets

2022-06-16 · Jean-Marc Valin, Ritwik Giri, Shrikant Venkataramani, Umut Isik 외

In real life, room effect, also known as room reverberation, and the present background noise degrade the quality of speech. Recently, deep learning-based speech enhancement approaches have shown a lot of promise and sur…

DenoisingSpeech Enhancement