paper-with-me

Papers

FLEXI: Benchmarking Full-duplex Human-LLM Speech Interaction

2025-09-26 · Yuan Ge, Saihan Chen, Jingqi Xiao, Xiaoqian Liu, Tong Xiao, Yan Xiang, Zhengtao Yu, Jingbo Zhu arxiv

Full-Duplex Speech-to-Speech Large Language Models (LLMs) are foundational to natural human-computer interaction, enabling real-time spoken dialogue systems. However, benchmarking and modeling these models remains a fundamental challenge. We introduce FLEXI, the first benchmark for full-duplex LLM-human spoken interaction that explicitly incorporates model interruption in emergency scenarios. FLEXI systematically evaluates the latency, quality, and conversational effectiveness of real-time dialogue through six diverse human-LLM interaction scenarios, revealing significant gaps between open source and commercial models in emergency awareness, turn terminating, and interaction latency. Finally, we suggest that next token-pair prediction offers a promising path toward achieving truly seamless and human-like full-duplex interaction.

📄 PDF Abstract BibTeX arXiv:2509.22243

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Enabling Conversational Behavior Reasoning Capabilities in Full-Duplex Speech

2025-12-25 · Shuchang Pan, Siddharth Banerjee, Dhruv Hebbar, Siddhant Patel 외 arxiv

Human conversation is organized by an implicit chain of thoughts that manifests as timed speech acts. Capturing this causal pathway is key to building natural full-duplex interactive systems. We introduce a framework tha…

Causal Inference

FlexDuo: A Pluggable System for Enabling Full-Duplex Capabilities in Speech Dialogue Systems

2025-02-19 · Borui Liao, Yulong Xu, Jiao Ou, Kaiyuan Yang 외

Full-Duplex Speech Dialogue Systems (Full-Duplex SDS) have significantly enhanced the naturalness of human-machine interaction by enabling real-time bidirectional communication. However, existing approaches face challeng…

Action DetectionActivity DetectionSpoken Dialogue Systems

FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems

2025-07-25 · Yizhou Peng, Yi-Wen Chao, Dianwen Ng, Yukun Ma 외 arxiv

Full-duplex spoken dialogue systems (FDSDS) enable more natural human-machine interactions by allowing real-time user interruptions and backchanneling, compared to traditional SDS that rely on turn-taking. However, exist…

DuplexChat: Constructing Speaker-Separated Full-Duplex Dialogue Speech at Scale for Spoken Dialogue Language Modeling

2026-07-06 · Wataru Nakata, Yuki Saito, Hiroshi Saruwatari arxiv

Full-duplex spoken dialogue models are trained on conversational speech in which each speaker is represented as a separate stream, but existing large-scale public speech corpora are mostly monaural, making them unsuited …

Speech Separation

Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency

2026-04-06 · Guan-Ting Lin, Chen Chen, Zhehuai Chen, Hung-yi Lee arxiv

We introduce Full-Duplex-Bench-v3 (FDB-v3), a benchmark for evaluating spoken language models under naturalistic speech conditions and multi-step tool use. Unlike prior work, our dataset consists entirely of real human a…