paper-with-me

Papers

Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities

2025-03-06 · Guan-Ting Lin, Jiachen Lian, Tingle Li, Qirui Wang, Gopala Anumanchipalli, Alexander H. Liu, Hung-Yi Lee

Spoken dialogue modeling poses challenges beyond text-based language modeling, requiring real-time interaction, turn-taking, and backchanneling. While most Spoken Dialogue Models (SDMs) operate in half-duplex mode-processing one turn at a time - emerging full-duplex SDMs can listen and speak simultaneously, enabling more natural conversations. However, current evaluations remain limited, focusing mainly on turn-based metrics or coarse corpus-level analyses. To address this, we introduce Full-Duplex-Bench, a benchmark that systematically evaluates key interactive behaviors: pause handling, backchanneling, turn-taking, and interruption management. Our framework uses automatic metrics for consistent, reproducible assessment and provides a fair, fast evaluation setup. By releasing our benchmark and code, we aim to advance spoken dialogue modeling and foster the development of more natural and engaging SDMs.

📄 PDF Abstract BibTeX arXiv:2503.04721

Code (1)

DanielLin94144/Full-Duplex-Bench 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingManagement

Similar Papers 제목 키워드 기반

VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents

2026-05-28 · Amrita Mazumdar, Seonwook Park, Rajarshi Roy, Nikhil Srihari 외 arxiv

Natural human conversation is full-duplex and audio-visual: people simultaneously speak and listen while continuously interpreting and producing nonverbal cues, such as nods, smiles, and gestures. To support successful h…

Visual Question Answering

FLEXI: Benchmarking Full-duplex Human-LLM Speech Interaction

2025-09-26 · Yuan Ge, Saihan Chen, Jingqi Xiao, Xiaoqian Liu 외 arxiv

Full-Duplex Speech-to-Speech Large Language Models (LLMs) are foundational to natural human-computer interaction, enabling real-time spoken dialogue systems. However, benchmarking and modeling these models remains a fund…

EchoChain: A Full-Duplex Benchmark for State-Update Reasoning Under Interruptions

2026-04-08 · Smit Nautambhai Modi, Gandharv Mahajan, Marc Wetter, Randall Welles arxiv

Real-time voice assistants must revise task state when users interrupt mid-response, but existing spoken-dialog benchmarks largely evaluate turn-based interaction and miss this failure mode. We introduce EchoChain, a con…

ECHO: A Matched-Contrast Benchmark for Context-Sensitive Turn-Taking in Full-Duplex Dialogue

2026-09-15 · Shuofeng Zhao, Hongwei Cai, Wenke Fan, Qingxiang Guo 외 arxiv

Full-duplex spoken dialogue systems must distinguish interruptions that require yielding the floor from backchannels that permit continued speaking. Existing benchmarks typically evaluate events independently and may the…

MTR-DuplexBench: Towards a Comprehensive Evaluation of Multi-Round Conversations for Full-Duplex Speech Language Models

2025-11-13 · He Zhang, Wenqian Cui, Haoning Xu, Xiaohui Li 외 arxiv

Full-Duplex Speech Language Models (FD-SLMs) enable real-time, overlapping conversational interactions, offering a more dynamic user experience compared to traditional half-duplex models. However, existing benchmarks pri…

Instruction Following