paper-with-me

Papers

C3: A Bilingual Benchmark for Spoken Dialogue Models Exploring Challenges in Complex Conversations

2025-07-30 · Chengqian Ma, Wei Tao, Yiwen Guo arxiv

Spoken Dialogue Models (SDMs) have recently attracted significant attention for their ability to generate voice responses directly to users' spoken queries. Despite their increasing popularity, there exists a gap in research focused on comprehensively understanding their practical effectiveness in comprehending and emulating human conversations. This is especially true compared to text-based Large Language Models (LLMs), which benefit from extensive benchmarking. Human voice interactions are inherently more complex than text due to characteristics unique to spoken dialogue. Ambiguity poses one challenge, stemming from semantic factors like polysemy, as well as phonological aspects such as heterograph, heteronyms, and stress patterns. Additionally, context-dependency, like omission, coreference, and multi-turn interaction, adds further complexity to human conversational dynamics. To illuminate the current state of SDM development and to address these challenges, we present a benchmark dataset in this paper, which comprises 1,079 instances in English and Chinese. Accompanied by an LLM-based evaluation method that closely aligns with human judgment, this dataset facilitates a comprehensive exploration of the performance of SDMs in tackling these practical challenges.

📄 PDF Abstract BibTeX arXiv:2507.22968

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SpokenWOZ: A Large-Scale Speech-Text Benchmark for Spoken Task-Oriented Dialogue Agents

2023-05-22 · NeurIPS 2023 11 · Shuzheng Si, Wentao Ma, Haoyu Gao, Yuchuan Wu 외

Task-oriented dialogue (TOD) models have made significant progress in recent years. However, previous studies primarily focus on datasets written by annotators, which has resulted in a gap between academic research and r…

Lost in Speech: Benchmarking, Evaluation, and Parsing of Spoken Bilingual Conversational Language Beyond Standard UD Assumptions

2026-02-06 · Nemika Tyagi, Olga Kellert, Holly Hendrix, Nelvin Licona-Guevara 외 arxiv

Spoken bilingual conversations pose substantial challenges for syntactic parsing because they often include disfluencies and discourse-driven structures that complicate dependency parsing under standard Universal Depende…

Dependency Parsing

Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities

2025-03-06 · Guan-Ting Lin, Jiachen Lian, Tingle Li, Qirui Wang 외

Spoken dialogue modeling poses challenges beyond text-based language modeling, requiring real-time interaction, turn-taking, and backchanneling. While most Spoken Dialogue Models (SDMs) operate in half-duplex mode-proces…

Language ModelingLanguage ModellingManagement

EVI: Multilingual Spoken Dialogue Tasks and Dataset for Knowledge-Based Enrolment, Verification, and Identification

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Knowledge-based authentication is crucial for task-oriented spoken dialogue systems that offer personalised and privacy-focused services. Such systems should be able to enrol (E), verify (V), and identify (I) new and rec…

Spoken Dialogue Systems

EVI: Multilingual Spoken Dialogue Tasks and Dataset for Knowledge-Based Enrolment, Verification, and Identification

2022-04-28 · Findings (NAACL) 2022 7 · Georgios P. Spithourakis, Ivan Vulić, Michał Lis, Iñigo Casanueva 외

Knowledge-based authentication is crucial for task-oriented spoken dialogue systems that offer personalised and privacy-focused services. Such systems should be able to enrol (E), verify (V), and identify (I) new and rec…

Speaker IdentificationSpeaker VerificationSpoken Dialogue Systems