Audio-visual Speaker Recognition with a Cross-modal Discriminative Network
Audio-visual speaker recognition is one of the tasks in the recent 2019 NIST speaker recognition evaluation (SRE). Studies in neuroscience and computer science all point to the fact that vision and auditory neural signals interact in the cognitive process. This motivated us to study a cross-modal network, namely voice-face discriminative network (VFNet) that establishes the general relation between human voice and face. Experiments show that VFNet provides additional speaker discriminative information. With VFNet, we achieve 16.54% equal error rate relative reduction over the score level fusion audio-visual baseline on evaluation set of 2019 NIST SRE.
Code (0)
등록된 구현이 없습니다.
Tasks
Speaker RecognitionSimilar Papers 제목 키워드 기반
Speaker Recognition in Realistic Scenario Using Multimodal Data
In recent years, an association is established between faces and voices of celebrities leveraging large scale audio-visual information from YouTube. The availability of large scale audio-visual datasets is instrumental i…
Speaker Recognition3D Convolutional Neural Networks for Cross Audio-Visual Matching Recognition
Audio-visual recognition (AVR) has been considered as a solution for speech recognition tasks when the audio is corrupted, as well as a visual recognition method used for speaker verification in multi-speaker scenarios. …
Speaker Verificationspeech-recognitionSpeech RecognitionTowards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition
Although many previous studies have carried out multimodal learning with real-time MRI data that captures the audio-visual kinematics of the vocal tract during speech, these studies have been limited by their reliance on…
Phoneme RecognitionThe 2021 NIST Speaker Recognition Evaluation
The 2021 Speaker Recognition Evaluation (SRE21) was the latest cycle of the ongoing evaluation series conducted by the U.S. National Institute of Standards and Technology (NIST) since 1996. It was the second large-scale …
Data AugmentationFace RecognitionPerson RecognitionSpeaker RecognitionOLKAVS: An Open Large-Scale Korean Audio-Visual Speech Dataset
Inspired by humans comprehending speech in a multi-modal manner, various audio-visual datasets have been constructed. However, most existing datasets focus on English, induce dependencies with various prediction models d…
Audio-Visual Speech RecognitionLip ReadingSpeaker Recognitionspeech-recognition+2