paper-with-me

Papers

Visual Cues Enhance Predictive Turn-Taking for Two-Party Human Interaction

2025-05-27 · Sam O'Connor Russell, Naomi Harte

Turn-taking is richly multimodal. Predictive turn-taking models (PTTMs) facilitate naturalistic human-robot interaction, yet most rely solely on speech. We introduce MM-VAP, a multimodal PTTM which combines speech with visual cues including facial expression, head pose and gaze. We find that it outperforms the state-of-the-art audio-only in videoconferencing interactions (84% vs. 79% hold/shift prediction accuracy). Unlike prior work which aggregates all holds and shifts, we group by duration of silence between turns. This reveals that through the inclusion of visual features, MM-VAP outperforms a state-of-the-art audio-only turn-taking model across all durations of speaker transitions. We conduct a detailed ablation study, which reveals that facial expression features contribute the most to model performance. Thus, our working hypothesis is that when interlocutors can see one another, visual cues are vital for turn-taking and must therefore be included for accurate turn-taking prediction. We additionally validate the suitability of automatic speech alignment for PTTM training using telephone speech. This work represents the first comprehensive analysis of multimodal PTTMs. We discuss implications for future work and make all code publicly available.

📄 PDF Abstract BibTeX arXiv:2505.21043

Code (1)

russelsa/mm-vap 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Visual Cues Support Robust Turn-taking Prediction in Noise

2025-05-28 · Sam O'Connor Russell, Naomi Harte

Accurate predictive turn-taking models (PTTMs) are essential for naturalistic human-robot interaction. However, little is known about their performance in noise. This study therefore explores PTTM performance in types of…

The Role of Prosodic and Lexical Cues in Turn-Taking with Self-Supervised Speech Representations

2026-01-20 · Sam OConnor Russell, Delphine Charuau, Naomi Harte arxiv

Fluid turn-taking remains a key challenge in human-robot interaction. Self-supervised speech representations (S3Rs) have driven many advances, but it remains unclear whether S3R-based turn-taking models rely on prosodic …

Toward Signing Activity Projection in Sign Language Interaction

2026-06-08 · Takao Obi, Wang Yusong, Koji Inoue, Kotaro Funakoshi arxiv

Social robots must interact robustly not only with users assumed by speech-centered systems but also with diverse users whose communication relies on different modalities, e.g., sign language. One important capability ga…

AV-Dialog: Spoken Dialogue Models with Audio-Visual Input

2025-11-14 · Tuochao Chen, Bandhav Veluri, Hongyu Gong, Shyamnath Gollakota arxiv

Dialogue models falter in noisy, multi-speaker environments, often producing irrelevant responses and awkward turn-taking. We present AV-Dialog, the first multimodal dialog framework that uses both audio and visual cues …

Boundary Detection

Automatic Evaluation of Turn-taking Cues in Conversational Speech Synthesis

2023-05-29 · Erik Ekstedt, Siyang Wang, Éva Székely, Joakim Gustafson 외

Turn-taking is a fundamental aspect of human communication where speakers convey their intention to either hold, or yield, their turn through prosodic cues. Using the recently proposed Voice Activity Projection model, we…

Speech Synthesistext-to-speechText to Speech