paper-with-me

Active Speaker Detection

1개 벤치마크 · 논문 66편 · 이 태스크의 논문 보기 →

Benchmarks

LRS3-TED

결과 1개

Most implemented

End-to-End Active Speaker Detection

2022-03-27 · 구현 3개

Papers

$C^3$ASD: Multi-Level Consistency-Driven Representation Learning

2026-07-03 · Jin Hong, Jisoo Park, Junseok Kwon arxiv

Active Speaker Detection determines whether a visible person in a video is speaking at each moment. While recent audio-visual fusion methods perform well on clean data, they degrade under real-world corruptions such as b…

Active Speaker DetectionRepresentation LearningKnowledge DistillationContrastive Learning

GateFusion: Hierarchical Gated Cross-Modal Fusion for Active Speaker Detection

2025-12-17 · Yu Wang, Juhyung Ha, Frangil M. Ramirez, Yuchen Wang 외 arxiv

Active Speaker Detection (ASD) aims to identify who is currently speaking in each frame of a video. Most state-of-the-art approaches rely on late fusion to combine visual and audio features, but late fusion often fails t…

Active Speaker Detection

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations

2025-07-29 · T. Aleksandra Ma, Sile Yin, Li-Chia Yang, Shuo Zhang arxiv

Speech enhancement in audio-only settings remains challenging, particularly in the presence of interfering speakers. This paper presents a simple yet effective real-time audio-visual speech enhancement (AVSE) system, RAV…

Audio-Visual Speech RecognitionActive Speaker DetectionSpeech Enhancement

UniTalk: Towards Universal Active Speaker Detection in Real World Scenarios

2025-05-28 · Le Thien Phuc Nguyen, Zhuoran Yu, Khoa Quang Nhat Cao, Yuwei Guo 외

We present UniTalk, a novel dataset specifically designed for the task of active speaker detection, emphasizing challenging scenarios to enhance model generalization. Unlike previously established benchmarks such as AVA,…

Active Speaker Detection

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization

2025-05-06 · Detao Bai, Zhiheng Ma, Xihan Wei, Liefeng Bo

The inherent synchronization between a speaker's lip movements, voice, and the underlying linguistic content offers a rich source of information for improving speech processing tasks, especially in challenging conditions…

Active Speaker DetectionAudio-Visual Speech RecognitionAudio-Visual SynchronizationRepresentation Learning+4

Understanding Co-speech Gestures in-the-wild

2025-03-28 · Sindhu B Hegde, K R Prajwal, Taein Kwon, Andrew Zisserman

Co-speech gestures play a vital role in non-verbal communication. In this paper, we introduce a new framework for co-speech gesture understanding in the wild. Specifically, we propose three new tasks and benchmarks to ev…

Active Speaker Detection

전체 66편 보기 →