paper-with-me

홈 › Papers

Audio-visual Speaker Recognition with a Cross-modal Discriminative Network

2020-08-10

Audio-visual speaker recognition is one of the tasks in the recent 2019 NIST speaker recognition evaluation (SRE). Studies in neuroscience and computer science all point to the fact that vision and auditory neural signals interact in the cognitive process. This motivated us to study a cross-modal network, namely voice-face discriminative network (VFNet) that establishes the general relation between human voice and face. Experiments show that VFNet provides additional speaker discriminative information. With VFNet, we achieve 16.54% equal error rate relative reduction over the score level fusion audio-visual baseline on evaluation set of 2019 NIST SRE.

📄 PDF Abstract BibTeX arXiv:2008.03894

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Recognition

Similar Papers 제목 키워드 기반

Speaker Recognition in Realistic Scenario Using Multimodal Data

2023-02-25 · Saqlain Hussain Shah, Muhammad Saad Saeed, Shah Nawaz, Muhammad Haroon Yousaf

In recent years, an association is established between faces and voices of celebrities leveraging large scale audio-visual information from YouTube. The availability of large scale audio-visual datasets is instrumental i…

Speaker Recognition

3D Convolutional Neural Networks for Cross Audio-Visual Matching Recognition

2017-06-18 · Amirsina Torfi, Seyed Mehdi Iranmanesh, Nasser M. Nasrabadi, Jeremy Dawson

Audio-visual recognition (AVR) has been considered as a solution for speech recognition tasks when the audio is corrupted, as well as a visual recognition method used for speaker verification in multi-speaker scenarios. …

Speaker Verificationspeech-recognitionSpeech Recognition

Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition

2025-05-29 · Sean Foley, Hong Nguyen, JIhwan Lee, Sudarsana Reddy Kadiri 외

Although many previous studies have carried out multimodal learning with real-time MRI data that captures the audio-visual kinematics of the vocal tract during speech, these studies have been limited by their reliance on…

Phoneme Recognition

The 2021 NIST Speaker Recognition Evaluation

2022-04-21 · Seyed Omid Sadjadi, Craig Greenberg, Elliot Singer, Lisa Mason 외

The 2021 Speaker Recognition Evaluation (SRE21) was the latest cycle of the ongoing evaluation series conducted by the U.S. National Institute of Standards and Technology (NIST) since 1996. It was the second large-scale …

Data AugmentationFace RecognitionPerson RecognitionSpeaker Recognition

OLKAVS: An Open Large-Scale Korean Audio-Visual Speech Dataset

2023-01-16 · Jeongkyun Park, Jung-Wook Hwang, Kwanghee Choi, Seung-Hyun Lee 외

Inspired by humans comprehending speech in a multi-modal manner, various audio-visual datasets have been constructed. However, most existing datasets focus on English, induce dependencies with various prediction models d…

Audio-Visual Speech RecognitionLip ReadingSpeaker Recognitionspeech-recognition+2