Audio-Visual Active Speaker Detection
2개 벤치마크 · 논문 25편 · 이 태스크의 논문 보기 →
Benchmarks
AVA-ActiveSpeaker
VPCD
Most implemented
Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection
End-to-End Active Speaker Detection
LoCoNet: Long-Short Context Network for Active Speaker Detection
Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection
AVA-ActiveSpeaker: An Audio-Visual Dataset for Active Speaker Detection
LASER: Lip Landmark Assisted Speaker Detection for Robustness
Papers
LASER: Lip Landmark Assisted Speaker Detection for Robustness
Active Speaker Detection (ASD) aims to identify speaking individuals in complex visual scenes. While humans can easily detect speech by matching lip movements to audio, current ASD models struggle to establish this corre…
Active Speaker DetectionAudio-Visual Active Speaker DetectionAn Efficient and Streaming Audio Visual Active Speaker Detection System
This paper delves into the challenging task of Active Speaker Detection (ASD), where the system needs to determine in real-time whether a person is speaking or not in a series of video frames. While previous works have m…
Active Speaker DetectionAudio-Visual Active Speaker DetectionCPUEnhancing Real-World Active Speaker Detection with Multi-Modal Extraction Pre-Training
Audio-visual active speaker detection (AV-ASD) aims to identify which visible face is speaking in a scene with one or more persons. Most existing AV-ASD methods prioritize capturing speech-lip correspondence. However, th…
Active Speaker DetectionAudio-Visual Active Speaker DetectionDenoisingTarget Speaker ExtractionTalkNCE: Improving Active Speaker Detection with Talk-Aware Contrastive Learning
The goal of this work is Active Speaker Detection (ASD), a task to determine whether a person is speaking or not in a series of video frames. Previous works have dealt with the task by exploring network architectures whi…
Active Speaker DetectionAudio-Visual Active Speaker DetectionContrastive LearningA Light Weight Model for Active Speaker Detection
Active speaker detection is a challenging task in audio-visual scenario understanding, which aims to detect who is speaking in one or more speakers scenarios. This task has received extensive attention as it is crucial i…
Active Speaker DetectionAudio-Visual Active Speaker Detectionmodelspeaker-diarization+2LoCoNet: Long-Short Context Network for Active Speaker Detection
Active Speaker Detection (ASD) aims to identify who is speaking in each frame of a video. ASD reasons from audio and visual information from two contexts: long-term intra-speaker context and short-term inter-speaker cont…
Active Speaker DetectionAudio-Visual Active Speaker Detection