paper-with-me

Papers

A Deep Neural Network for Short-Segment Speaker Recognition

2019-07-22 · Amirhossein Hajavi, Ali Etemad

Todays interactive devices such as smart-phone assistants and smart speakers often deal with short-duration speech segments. As a result, speaker recognition systems integrated into such devices will be much better suited with models capable of performing the recognition task with short-duration utterances. In this paper, a new deep neural network, UtterIdNet, capable of performing speaker recognition with short speech segments is proposed. Our proposed model utilizes a novel architecture that makes it suitable for short-segment speaker recognition through an efficiently increased use of information in short speech segments. UtterIdNet has been trained and tested on the VoxCeleb datasets, the latest benchmarks in speaker recognition. Evaluations for different segment durations show consistent and stable performance for short segments, with significant improvement over the previous models for segments of 2 seconds, 1 second, and especially sub-second durations (250 ms and 500 ms).

📄 PDF Abstract BibTeX arXiv:1907.10420

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Recognition

Similar Papers 제목 키워드 기반

Universal speaker recognition encoders for different speech segments duration

2022-10-28 · Sergey Novoselov, Vladimir Volokhov, Galina Lavrentyeva

Creating universal speaker encoders which are robust for different acoustic and speech duration conditions is a big challenge today. According to our observations systems trained on short speech segments are optimal for …

Speaker RecognitionSpeaker Verification

Self Multi-Head Attention for Speaker Recognition

2019-06-24 · Miquel India, Pooyan Safari, Javier Hernando

Most state-of-the-art Deep Learning (DL) approaches for speaker recognition work on a short utterance level. Given the speech signal, these algorithms extract a sequence of speaker embeddings from short segments and thos…

Speaker Recognition

Character-aware audio-visual subtitling in context

2024-10-14 · Jaesung Huh, Andrew Zisserman

This paper presents an improved framework for character-aware audio-visual subtitling in TV shows. Our approach integrates speech recognition, speaker diarisation, and character recognition, utilising both audio and visu…

Language ModellingLarge Language Modelspeech-recognitionSpeech Recognition

Neural Predictive Coding using Convolutional Neural Networks towards Unsupervised Learning of Speaker Characteristics

2018-02-22 · Arindam Jati, Panayiotis Georgiou

Learning speaker-specific features is vital in many applications like speaker recognition, diarization and speech recognition. This paper provides a novel approach, we term Neural Predictive Coding (NPC), to learn speake…

Speaker IdentificationSpeaker RecognitionSpeaker Verificationspeech-recognition+1

Novel Quality Metric for Duration Variability Compensation in Speaker Verification using i-Vectors

2018-12-03 · Arnab Poddar, Md Sahidullah, Goutam Saha

Automatic speaker verification (ASV) is the process to recognize persons using voice as biometric. The ASV systems show considerable recognition performance with sufficient amount of speech from matched condition. One of…

Speaker Verification