paper-with-me

Papers

Visual Cues Support Robust Turn-taking Prediction in Noise

2025-05-28 · Sam O'Connor Russell, Naomi Harte

Accurate predictive turn-taking models (PTTMs) are essential for naturalistic human-robot interaction. However, little is known about their performance in noise. This study therefore explores PTTM performance in types of noise likely to be encountered once deployed. Our analyses reveal PTTMs are highly sensitive to noise. Hold/shift accuracy drops from 84% in clean speech to just 52% in 10 dB music noise. Training with noisy data enables a multimodal PTTM, which includes visual features to better exploit visual cues, with 72% accuracy in 10 dB music noise. The multimodal PTTM outperforms the audio-only PTTM across all noise types and SNRs, highlighting its ability to exploit visual cues; however, this does not always generalise to new types of noise. Analysis also reveals that successful training relies on accurate transcription, limiting the use of ASR-derived transcriptions to clean conditions. We make code publicly available for future research.

📄 PDF Abstract BibTeX arXiv:2505.22088

Code (1)

russelsa/mm-vap 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Visual Cues Enhance Predictive Turn-Taking for Two-Party Human Interaction

2025-05-27 · Sam O'Connor Russell, Naomi Harte

Turn-taking is richly multimodal. Predictive turn-taking models (PTTMs) facilitate naturalistic human-robot interaction, yet most rely solely on speech. We introduce MM-VAP, a multimodal PTTM which combines speech with v…

The Role of Prosodic and Lexical Cues in Turn-Taking with Self-Supervised Speech Representations

2026-01-20 · Sam OConnor Russell, Delphine Charuau, Naomi Harte arxiv

Fluid turn-taking remains a key challenge in human-robot interaction. Self-supervised speech representations (S3Rs) have driven many advances, but it remains unclear whether S3R-based turn-taking models rely on prosodic …

Multimodal Continuous Turn-Taking Prediction Using Multiscale RNNs

2018-08-31 · Matthew Roddy, Gabriel Skantze, Naomi Harte

In human conversational interactions, turn-taking exchanges can be coordinated using cues from multiple modalities. To design spoken dialog systems that can conduct fluid interactions it is desirable to incorporate cues …

Prediction

AV-Dialog: Spoken Dialogue Models with Audio-Visual Input

2025-11-14 · Tuochao Chen, Bandhav Veluri, Hongyu Gong, Shyamnath Gollakota arxiv

Dialogue models falter in noisy, multi-speaker environments, often producing irrelevant responses and awkward turn-taking. We present AV-Dialog, the first multimodal dialog framework that uses both audio and visual cues …

Boundary Detection

Modeling Turn-Taking with Semantically Informed Gestures

2025-10-22 · Varsha Suresh, M. Hamza Mughal, Christian Theobalt, Vera Demberg arxiv

In conversation, humans use multimodal cues, such as speech, gestures, and gaze, to manage turn-taking. While linguistic and acoustic features are informative, gestures provide complementary cues for modeling these trans…