Active Speaker Detection
1개 벤치마크 · 논문 66편 · 이 태스크의 논문 보기 →
Benchmarks
LRS3-TED
Most implemented
Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection
End-to-End Active Speaker Detection
LoCoNet: Long-Short Context Network for Active Speaker Detection
Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection
AVA-ActiveSpeaker: An Audio-Visual Dataset for Active Speaker Detection
UniTalk: Towards Universal Active Speaker Detection in Real World Scenarios
Papers
$C^3$ASD: Multi-Level Consistency-Driven Representation Learning
Active Speaker Detection determines whether a visible person in a video is speaking at each moment. While recent audio-visual fusion methods perform well on clean data, they degrade under real-world corruptions such as b…
Active Speaker DetectionRepresentation LearningKnowledge DistillationContrastive LearningGateFusion: Hierarchical Gated Cross-Modal Fusion for Active Speaker Detection
Active Speaker Detection (ASD) aims to identify who is currently speaking in each frame of a video. Most state-of-the-art approaches rely on late fusion to combine visual and audio features, but late fusion often fails t…
Active Speaker DetectionReal-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations
Speech enhancement in audio-only settings remains challenging, particularly in the presence of interfering speakers. This paper presents a simple yet effective real-time audio-visual speech enhancement (AVSE) system, RAV…
Audio-Visual Speech RecognitionActive Speaker DetectionSpeech EnhancementUniTalk: Towards Universal Active Speaker Detection in Real World Scenarios
We present UniTalk, a novel dataset specifically designed for the task of active speaker detection, emphasizing challenging scenarios to enhance model generalization. Unlike previously established benchmarks such as AVA,…
Active Speaker DetectionCoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization
The inherent synchronization between a speaker's lip movements, voice, and the underlying linguistic content offers a rich source of information for improving speech processing tasks, especially in challenging conditions…
Active Speaker DetectionAudio-Visual Speech RecognitionAudio-Visual SynchronizationRepresentation Learning+4Understanding Co-speech Gestures in-the-wild
Co-speech gestures play a vital role in non-verbal communication. In this paper, we introduce a new framework for co-speech gesture understanding in the wild. Specifically, we propose three new tasks and benchmarks to ev…
Active Speaker Detection