Multi-Task Learning for Audio Visual Active Speaker Detection
This report describes the approach underlying our submission to the active speaker detection task (task B-2) of ActivityNet Challenge 2019. We introduce a new audio-visual model which builds upon a 3D-ResNet18 visual model pretrained for lipreading and a VGG-M acoustic model pretrained for audio-to-video synchronization. The model is trained with two losses in a multi-task learning fashion: a contrastive loss to enforce matching between audio and video features for active speakers, and a regular crossentropy loss to obtain speaker / non-speaker labels. This model obtains 84.0% mAP on the validation set of AVAActiveSpeaker. Experimental results showcase the pretrained embeddings' abilities to transfer across tasks and data formats, as well as the advantage of the proposed multi-task learning strategy.
Code (0)
등록된 구현이 없습니다.
Tasks
Active Speaker DetectionAudio-Visual Active Speaker DetectionLipreadingMulti-Task LearningVideo SynchronizationSimilar Papers 제목 키워드 기반
Look\&Listen: Multi-Modal Correlation Learning for Active Speaker Detection and Speech Enhancement
Active speaker detection and speech enhancement have become two increasingly attractive topics in audio-visual scenario understanding. According to their respective characteristics, the scheme of independently designed a…
Active Speaker DetectionMulti-Task LearningSpeech EnhancementActive Speakers in Context
Current methods for active speak er detection focus on modeling short-term audiovisual information from a single speaker. Although this strategy can be enough for addressing single-speaker scenarios, it prevents accurate…
Active Speaker DetectionAudio-Visual Active Speaker DetectionAVA-ActiveSpeaker: An Audio-Visual Dataset for Active Speaker Detection
Active speaker detection is an important component in video analysis algorithms for applications such as speaker diarization, video re-targeting for meetings, speech enhancement, and human-robot interaction. The absence …
Active Speaker DetectionAudio-Visual Active Speaker DetectionDiversityspeaker-diarization+2Leveraging Visual Supervision for Array-based Active Speaker Detection and Localization
Conventional audio-visual approaches for active speaker detection (ASD) typically rely on visually pre-extracted face tracks and the corresponding single-channel audio to find the speaker in a video. Therefore, they tend…
Active Speaker DetectionSelf-Supervised LearningDeep Learning Based Audio-Visual Multi-Speaker DOA Estimation Using Permutation-Free Loss Function
In this paper, we propose a deep learning based multi-speaker direction of arrival (DOA) estimation with audio and visual signals by using permutation-free loss function. We first collect a data set for multi-modal sound…
Active Speaker DetectionSound Source Localization