paper-with-me

홈 › Papers

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

2026-07-16 · David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer, Jakub Piotr Cłapa, Theo Lebryk, Jens Madsen, Olya Ossipova, Sharath Rao, Hoon Shin, Tigran Soghbatyan, Georg Streich, Rashish Tandon, Panagiotis Tzirakis arxiv

Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation. To this end, we introduce the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI across text-to-speech (TTS), speech-to-speech (STS), speech understanding (SU), and automatic speech recognition (ASR). Our evaluations indicate that performance is highly dimension-specific. For TTS, naturalness, expressiveness, identity stability, and reliability are largely independent evaluation dimensions. For STS, access to audio does not guarantee use of vocal affect, and some agents remain largely transcript-driven. For SU, models perform unevenly across paralinguistic tasks. For ASR, real world accent, emotion, noise, and conversational conditions expose failures that are not captured by established clean-speech benchmarks. Together, these results show that voice AI should be evaluated as a profile of acoustic, expressive, interactional, and robustness capabilities rather than by a single aggregate score.

📄 PDF Abstract BibTeX arXiv:2607.14846

Code (3)

InsomaniacElf/sg-tamil-tts-resources- ★ 1
iszhanjiawei/TTS_arxiv_daily ★ 3
liutaocode/TTS-arxiv-daily ★ 660

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

$τ$-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains

2026-03-14 · Soham Ray, Keshav Dhandhania, Victor Barres, Karthik Narasimhan arxiv

Full-duplex voice agents--systems that listen and speak simultaneously--are rapidly moving from research to production. However, existing evaluations address conversational dynamics and task completion in isolation. We i…

VoiceWukong: Benchmarking Deepfake Voice Detection

2024-09-10 · Ziwei Yan, Yanjie Zhao, Haoyu Wang

With the rapid advancement of technologies like text-to-speech (TTS) and voice conversion (VC), detecting deepfake voices has become increasingly crucial. However, both academia and industry lack a comprehensive and intu…

BenchmarkingFace SwappingLarge Language Modeltext-to-speech+2

VoiceBench: Benchmarking LLM-Based Voice Assistants

2024-10-22 · Yiming Chen, Xianghu Yue, Chen Zhang, Xiaoxue Gao 외

Building on the success of large language models (LLMs), recent advancements such as GPT-4o have enabled real-time speech interactions through LLM-based voice assistants, offering a significantly improved user experience…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)BenchmarkingGeneral Knowledge+2

Back to Basics: Revisiting ASR in the Age of Voice Agents

2026-03-26 · Geeyang Tay, Wentao Ma, Jaewon Lee, Yuzhi Tang 외 arxiv

Automatic speech recognition (ASR) systems have achieved near-human accuracy on curated benchmarks, yet still fail in real-world voice agents under conditions that current evaluations do not systematically cover. Without…

Speech Recognition

EchoChain: A Full-Duplex Benchmark for State-Update Reasoning Under Interruptions

2026-04-08 · Smit Nautambhai Modi, Gandharv Mahajan, Marc Wetter, Randall Welles arxiv

Real-time voice assistants must revise task state when users interrupt mid-response, but existing spoken-dialog benchmarks largely evaluate turn-based interaction and miss this failure mode. We introduce EchoChain, a con…