A Multi-View Approach To Audio-Visual Speaker Verification
Although speaker verification has conventionally been an audio-only task, some practical applications provide both audio and visual streams of input. In these cases, the visual stream provides complementary information and can often be leveraged in conjunction with the acoustics of speech to improve verification performance. In this study, we explore audio-visual approaches to speaker verification, starting with standard fusion techniques to learn joint audio-visual (AV) embeddings, and then propose a novel approach to handle cross-modal verification at test time. Specifically, we investigate unimodal and concatenation based AV fusion and report the lowest AV equal error rate (EER) of 0.7% on the VoxCeleb1 dataset using our best system. As these methods lack the ability to do cross-modal verification, we introduce a multi-view model which uses a shared classifier to map audio and video into the same space. This new approach achieves 28% EER on VoxCeleb1 in the challenging testing condition of cross-modal verification.
Code (0)
등록된 구현이 없습니다.
Tasks
Speaker VerificationSimilar Papers 제목 키워드 기반
Getting More for Less: Using Weak Labels and AV-Mixup for Robust Audio-Visual Speaker Verification
Distance Metric Learning (DML) has typically dominated the audio-visual speaker verification problem space, owing to strong performance in new and unseen classes. In our work, we explored multitask learning techniques to…
Metric LearningMulti-Task LearningSpeaker VerificationAudio-Visual Speaker Verification via Joint Cross-Attention
Speaker verification has been widely explored using speech signals, which has shown significant improvement using deep models. Recently, there has been a surge in exploring faces and voices as they can offer more complem…
Speaker VerificationSSAVSV: Towards Unified Model for Self-Supervised Audio-Visual Speaker Verification
Conventional audio-visual methods for speaker verification rely on large amounts of labeled data and separate modality-specific architectures, which is computationally expensive, limiting their scalability. To address th…
Contrastive LearningSelf-Supervised LearningSpeaker VerificationMargin-Mixup: A Method for Robust Speaker Verification in Multi-Speaker Audio
This paper is concerned with the task of speaker verification on audio with multiple overlapping speakers. Most speaker verification systems are designed with the assumption of a single speaker being present in a given a…
Speaker VerificationLearning Lip-Based Audio-Visual Speaker Embeddings with AV-HuBERT
This paper investigates self-supervised pre-training for audio-visual speaker representation learning where a visual stream showing the speaker's mouth area is used alongside speech as inputs. Our study focuses on the Au…
Representation LearningSpeaker Verification