paper-with-me

Papers

Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics

2025-03-03 · Siddhant Arora, Zhiyun Lu, Chung-Cheng Chiu, Ruoming Pang, Shinji Watanabe

The recent wave of audio foundation models (FMs) could provide new capabilities for conversational modeling. However, there have been limited efforts to evaluate these audio FMs comprehensively on their ability to have natural and interactive conversations. To engage in meaningful conversation with the end user, we would want the FMs to additionally perform a fluent succession of turns without too much overlapping speech or long stretches of silence. Inspired by this, we ask whether the recently proposed audio FMs can understand, predict, and perform turn-taking events? To answer this, we propose a novel evaluation protocol that can assess spoken dialog system's turn-taking capabilities using a supervised model as a judge that has been trained to predict turn-taking events in human-human conversations. Using this protocol, we present the first comprehensive user study that evaluates existing spoken dialogue systems on their ability to perform turn-taking events and reveal many interesting insights, such as they sometimes do not understand when to speak up, can interrupt too aggressively and rarely backchannel. We further evaluate multiple open-source and proprietary audio FMs accessible through APIs on carefully curated test benchmarks from Switchboard to measure their ability to understand and predict turn-taking events and identify significant room for improvement. We will open source our evaluation platform to promote the development of advanced conversational AI systems.

📄 PDF Abstract BibTeX arXiv:2503.01174

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingSpoken Dialogue Systems

Similar Papers 제목 키워드 기반

Extreme-scale Talking-Face Video Upsampling with Audio-Visual Priors

2022-08-17 · Sindhu B Hegde, Rudrabha Mukhopadhyay, Vinay P Namboodiri, C. V. Jawahar

In this paper, we explore an interesting question of what can be obtained from an $8\times8$ pixel video sequence. Surprisingly, it turns out to be quite a lot. We show that when we process this $8\times8$ video with the…

Super-ResolutionVideo Compression

Deep Multi-modality Soft-decoding of Very Low Bit-rate Face Videos

2020-08-02 · Yanhui Guo, Xi Zhang, Xiaolin Wu

We propose a novel deep multi-modality neural network for restoring very low bit rate videos of talking heads. Such video contents are very common in social media, teleconferencing, distance education, tele-medicine, etc…

QuantizationVideo CompressionVideo Restoration

Decoupled Self-Forcing Distillation for Streaming Talking Head Generation

2026-09-09 · Yanru An, Ruiyan Wang, Wenwu Wei, Rui Bu 외 arxiv

Streaming talking-head generation produces each frame as its driving audio arrives, yet fidelity and efficiency have so far pulled in opposite directions: end-to-end methods condition a video diffusion model on audio dir…

Talking Head Generation

DiffTalk: Crafting Diffusion Models for Generalized Audio-Driven Portraits Animation

2023-01-10 · CVPR 2023 1 · Shuai Shen, Wenliang Zhao, Zibin Meng, Wanhua Li 외

Talking head synthesis is a promising approach for the video production industry. Recently, a lot of effort has been devoted in this research area to improve the generation quality or enhance the model generalization. Ho…

DenoisingTalking Head Generation

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

2025-06-03 · Chetwin Low, Weimin WANG

In this paper, we present TalkingMachines -- an efficient framework that transforms pretrained video generation models into real-time, audio-driven character animators. TalkingMachines enables natural conversational expe…

DecoderKnowledge DistillationLanguage ModelingLanguage Modelling+2