paper-with-me

홈 › Papers

VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context

2025-11-11 · Heyang Liu, Ziyang Cheng, Yuhao Wang, Hongcheng Liu, Yiqi Li, Ronghua Wu, Qunshan Gu, Yanfeng Wang, Yu Wang arxiv

The development of multi-modal large language models (LLMs) leads to intelligent approaches capable of speech interactions. As one of the most widely spoken languages globally, Mandarin is supported by most models to enhance their applicability and reach. However, the scarcity of comprehensive speech-to-speech (S2S) benchmarks in Mandarin contexts impedes systematic evaluation for developers and hinders fair model comparison for users. In this work, we propose VocalBench-zh, an ability-level divided evaluation suite adapted to Mandarin context consisting of 10 well-crafted subsets and over 10K high-quality instances, covering 12 user-oriented characters. The evaluation experiment on 14 mainstream models reveals the common challenges for current routes, and highlights the need for new insights into next-generation speech interactive systems. The evaluation codes and datasets will be available at https://github.com/SJTU-OmniAgent/VocalBench-zh.

📄 PDF Abstract BibTeX arXiv:2511.08230

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models

2025-05-21 · Heyang Liu, Yuhao Wang, Ziyang Cheng, Ronghua Wu 외

The rapid advancement of large language models (LLMs) has accelerated the development of multi-modal models capable of vocal communication. Unlike text-based interactions, speech conveys rich and diverse information, inc…

Benchmarking

VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency

2025-10-17 · Hongcheng Liu, Yixuan Hou, Heyang Liu, Yuhao Wang 외 arxiv

While Speech Large Language Models (Speech-LLMs) show strong performance in many applications, their robustness is critically under-tested, especially to speech disfluency. Existing evaluations often rely on idealized in…

Enabling Conversational Behavior Reasoning Capabilities in Full-Duplex Speech

2025-12-25 · Shuchang Pan, Siddharth Banerjee, Dhruv Hebbar, Siddhant Patel 외 arxiv

Human conversation is organized by an implicit chain of thoughts that manifests as timed speech acts. Capturing this causal pathway is key to building natural full-duplex interactive systems. We introduce a framework tha…

Causal Inference

WER We Stand: Benchmarking Urdu ASR Models

2024-09-17 · Samee Arif, Sualeha Farid, Aamina Jamal Khan, Mustafa Abbas 외

This paper presents a comprehensive evaluation of Urdu Automatic Speech Recognition (ASR) models. We analyze the performance of three ASR model families: Whisper, MMS, and Seamless-M4T using Word Error Rate (WER), along …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Benchmarkingspeech-recognition+2

ASR Benchmarking: Need for a More Representative Conversational Dataset

2024-09-18 · Gaurav Maheshwari, Dmitry Ivanov, Théo Johannet, Kevin El Haddad

Automatic Speech Recognition (ASR) systems have achieved remarkable performance on widely used benchmarks such as LibriSpeech and Fleurs. However, these benchmarks do not adequately reflect the complexities of real-world…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Benchmarkingspeech-recognition+1