Rethinking Audio-visual Synchronization for Active Speaker Detection
Active speaker detection (ASD) systems are important modules for analyzing multi-talker conversations. They aim to detect which speakers or none are talking in a visual scene at any given time. Existing research on ASD does not agree on the definition of active speakers. We clarify the definition in this work and require synchronization between the audio and visual speaking activities. This clarification of definition is motivated by our extensive experiments, through which we discover that existing ASD methods fail in modeling the audio-visual synchronization and often classify unsynchronized videos as active speaking. To address this problem, we propose a cross-modal contrastive learning strategy and apply positional encoding in attention modules for supervised ASD models to leverage the synchronization cue. Experimental results suggest that our model can successfully detect unsynchronized speaking as not speaking, addressing the limitation of current models.
Code (0)
등록된 구현이 없습니다.
Tasks
Active Speaker DetectionAudio-Visual SynchronizationContrastive LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Target Active Speaker Detection with Audio-visual Cues
In active speaker detection (ASD), we would like to detect whether an on-screen person is speaking based on audio-visual cues. Previous studies have primarily focused on modeling audio-visual synchronization cue, which d…
Active Speaker DetectionAudio-Visual SynchronizationMulti-Task Learning for Audio Visual Active Speaker Detection
This report describes the approach underlying our submission to the active speaker detection task (task B-2) of ActivityNet Challenge 2019. We introduce a new audio-visual model which builds upon a 3D-ResNet18 visual mod…
Active Speaker DetectionAudio-Visual Active Speaker DetectionLipreadingMulti-Task Learning+1CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization
The inherent synchronization between a speaker's lip movements, voice, and the underlying linguistic content offers a rich source of information for improving speech processing tasks, especially in challenging conditions…
Active Speaker DetectionAudio-Visual Speech RecognitionAudio-Visual SynchronizationRepresentation Learning+4UniSync: A Unified Framework for Audio-Visual Synchronization
Precise audio-visual synchronization in speech videos is crucial for content quality and viewer comprehension. Existing methods have made significant strides in addressing this challenge through rule-based approaches and…
Audio-Visual SynchronizationContrastive LearningFace GenerationFace Parsing+1LPIPS-AttnWav2Lip: Generic Audio-Driven lip synchronization for Talking Head Generation in the Wild
Researchers have shown a growing interest in Audio-driven Talking Head Generation. The primary challenge in talking head generation is achieving audio-visual coherence between the lips and the audio, known as lip synchro…
Talking Head Generation