paper-with-me

홈 › Papers

Speaker Recognition in Realistic Scenario Using Multimodal Data

2023-02-25 · Saqlain Hussain Shah, Muhammad Saad Saeed, Shah Nawaz, Muhammad Haroon Yousaf

In recent years, an association is established between faces and voices of celebrities leveraging large scale audio-visual information from YouTube. The availability of large scale audio-visual datasets is instrumental in developing speaker recognition methods based on standard Convolutional Neural Networks. Thus, the aim of this paper is to leverage large scale audio-visual information to improve speaker recognition task. To achieve this task, we proposed a two-branch network to learn joint representations of faces and voices in a multimodal system. Afterwards, features are extracted from the two-branch network to train a classifier for speaker recognition. We evaluated our proposed framework on a large scale audio-visual dataset named VoxCeleb$1$. Our results show that addition of facial information improved the performance of speaker recognition. Moreover, our results indicate that there is an overlap between face and voice.

📄 PDF Abstract BibTeX arXiv:2302.13033

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Recognition

Similar Papers 제목 키워드 기반

SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models

2025-08-08 · Han Yin, Yafeng Chen, Chong Deng, Luyao Cheng 외 arxiv

The Speaker Diarization and Recognition (SDR) task aims to predict "who spoke when and what" within an audio clip, which is a crucial task in various real-world multi-speaker scenarios such as meeting transcription and d…

Speaker DiarizationSpeech Recognition

Analysis of Deep Clustering as Preprocessing for Automatic Speech Recognition of Sparsely Overlapping Speech

2019-05-09 · Tobias Menne, Ilya Sklyar, Ralf Schlüter, Hermann Ney

Significant performance degradation of automatic speech recognition (ASR) systems is observed when the audio signal contains cross-talk. One of the recently proposed approaches to solve the problem of multi-speaker ASR i…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ClusteringDeep Clustering+2

End-to-End Single-Channel Speaker-Turn Aware Conversational Speech Translation

2023-11-01 · Juan Zuluaga-Gomez, Zhaocheng Huang, Xing Niu, Rohit Paturi 외

Conventional speech-to-text translation (ST) systems are trained on single-speaker utterances, and they may not generalize to real-life scenarios where the audio contains conversations by multiple speakers. In this paper…

Automatic Speech Recognitionspeech-recognitionSpeech RecognitionSpeech-to-Text+2

Channel adversarial training for cross-channel text-independent speaker recognition

2019-02-25

The conventional speaker recognition frameworks (e.g., the i-vector and CNN-based approach) have been successfully applied to various tasks when the channel of the enrolment dataset is similar to that of the test dataset…

Domain AdaptationSpeaker RecognitionText-Independent Speaker Recognition

DeepMSRF: A novel Deep Multimodal Speaker Recognition framework with Feature selection

2020-07-14 · Ehsan Asali, Farzan Shenavarmasouleh, Farid Ghareh Mohammadi, Prasanth Sengadu Suresh 외

For recognizing speakers in video streams, significant research studies have been made to obtain a rich machine learning model by extracting high-level speaker's features such as facial expression, emotion, and gender. H…

feature selectionSpeaker Recognition