paper-with-me

Papers

Rethinking Audio-visual Synchronization for Active Speaker Detection

2022-06-21 · Abudukelimu Wuerkaixi, You Zhang, Zhiyao Duan, ChangShui Zhang

Active speaker detection (ASD) systems are important modules for analyzing multi-talker conversations. They aim to detect which speakers or none are talking in a visual scene at any given time. Existing research on ASD does not agree on the definition of active speakers. We clarify the definition in this work and require synchronization between the audio and visual speaking activities. This clarification of definition is motivated by our extensive experiments, through which we discover that existing ASD methods fail in modeling the audio-visual synchronization and often classify unsynchronized videos as active speaking. To address this problem, we propose a cross-modal contrastive learning strategy and apply positional encoding in attention modules for supervised ASD models to leverage the synchronization cue. Experimental results suggest that our model can successfully detect unsynchronized speaking as not speaking, addressing the limitation of current models.

📄 PDF Abstract BibTeX arXiv:2206.10421

Code (0)

등록된 구현이 없습니다.

Tasks

Active Speaker DetectionAudio-Visual SynchronizationContrastive Learning

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Target Active Speaker Detection with Audio-visual Cues

2023-05-22 · Yidi Jiang, Ruijie Tao, Zexu Pan, Haizhou Li

In active speaker detection (ASD), we would like to detect whether an on-screen person is speaking based on audio-visual cues. Previous studies have primarily focused on modeling audio-visual synchronization cue, which d…

Active Speaker DetectionAudio-Visual Synchronization

Multi-Task Learning for Audio Visual Active Speaker Detection

2019-06-01 · The ActivityNet Large-Scale Activity Recognition Challenge Workshop, CVPR 2019 6 · Yuanhang Zhang, Jingyun Xiao, Shuang Yang, Shiguang Shan

This report describes the approach underlying our submission to the active speaker detection task (task B-2) of ActivityNet Challenge 2019. We introduce a new audio-visual model which builds upon a 3D-ResNet18 visual mod…

Active Speaker DetectionAudio-Visual Active Speaker DetectionLipreadingMulti-Task Learning+1

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization

2025-05-06 · Detao Bai, Zhiheng Ma, Xihan Wei, Liefeng Bo

The inherent synchronization between a speaker's lip movements, voice, and the underlying linguistic content offers a rich source of information for improving speech processing tasks, especially in challenging conditions…

Active Speaker DetectionAudio-Visual Speech RecognitionAudio-Visual SynchronizationRepresentation Learning+4

UniSync: A Unified Framework for Audio-Visual Synchronization

2025-03-20 · Tao Feng, Yifan Xie, Xun Guan, Jiyuan Song 외

Precise audio-visual synchronization in speech videos is crucial for content quality and viewer comprehension. Existing methods have made significant strides in addressing this challenge through rule-based approaches and…

Audio-Visual SynchronizationContrastive LearningFace GenerationFace Parsing+1

LPIPS-AttnWav2Lip: Generic Audio-Driven lip synchronization for Talking Head Generation in the Wild

2026-01-30 · Zhipeng Chen, Xinheng Wang, Lun Xie, Haijie Yuan 외 arxiv

Researchers have shown a growing interest in Audio-driven Talking Head Generation. The primary challenge in talking head generation is achieving audio-visual coherence between the lips and the audio, known as lip synchro…

Talking Head Generation