paper-with-me

홈 › Papers

Speaker Recognition with Random Digit Strings Using Uncertainty Normalized HMM-based i-vectors

2019-07-13 · Nooshin Maghsoodi, Hossein Sameti, Hossein Zeinali, Themos~Stafylakis

In this paper, we combine Hidden Markov Models (HMMs) with i-vector extractors to address the problem of text-dependent speaker recognition with random digit strings. We employ digit-specific HMMs to segment the utterances into digits, to perform frame alignment to HMM states and to extract Baum-Welch statistics. By making use of the natural partition of input features into digits, we train digit-specific i-vector extractors on top of each HMM and we extract well-localized i-vectors, each modelling merely the phonetic content corresponding to a single digit. We then examine ways to perform channel and uncertainty compensation, and we propose a novel method for using the uncertainty in the i-vector estimates. The experiments on RSR2015 part III show that the proposed method attains 1.52\% and 1.77\% Equal Error Rate (EER) for male and female respectively, outperforming state-of-the-art methods such as x-vectors, trained on vast amounts of data. Furthermore, these results are attained by a single system trained entirely on RSR2015, and by a simple score-normalized cosine distance. Moreover, we show that the omission of channel compensation yields only a minor degradation in performance, meaning that the system attains state-of-the-art results even without recordings from multiple handsets per speaker for training or enrolment. Similar conclusions are drawn from our experiments on the RedDots corpus, where the same method is evaluated on phrases. Finally, we report results with bottleneck features and show that further improvement is attained when fusing them with spectral features.

📄 PDF Abstract BibTeX arXiv:1907.06111

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker RecognitionSpeaker Verification

Similar Papers 제목 키워드 기반

Revisiting joint decoding based multi-talker speech recognition with DNN acoustic model

2021-10-31 · Martin Kocour, Kateřina Žmolíková, Lucas Ondel, Ján Švec 외

In typical multi-talker speech recognition systems, a neural network-based acoustic model predicts senone state posteriors for each speaker. These are later used by a single-talker decoder which is applied on each speake…

Decoderspeech-recognitionSpeech Recognition

End-to-End Approach for Recognition of Historical Digit Strings

2021-04-28 · Mengqiao Zhao, Andre G. Hochuli, Abbas Cheddad

The plethora of digitalised historical document datasets released in recent years has rekindled interest in advancing the field of handwriting pattern recognition. In the same vein, a recently published data set, known a…

Handwriting RecognitionSegmentation

End-to-end multi-talker audio-visual ASR using an active speaker attention module

2022-04-01 · Richard Rose, Olivier Siohan

This paper presents a new approach for end-to-end audio-visual multi-talker speech recognition. The approach, referred to here as the visual context attention model (VCAM), is important because it uses the available vide…

speech-recognitionSpeech Recognition

Segmentation-Free Approaches for Handwritten Numeral String Recognition

2018-04-24 · Andre G. Hochuli, Luiz E. S. Oliveira, Alceu S. Britto Jr, Robert Sabourin

This paper presents segmentation-free strategies for the recognition of handwritten numeral strings of unknown length. A synthetic dataset of touching numeral strings of sizes 2-, 3- and 4-digits was created to train end…

Segmentation

Implementation Of Back-Propagation Neural Network For Isolated Bangla Speech Recognition

2013-08-17 · Md. Ali Hossain, Md. Mijanur Rahman, Uzzal Kumar Prodhan, Md. Farukuzzaman Khan

This paper is concerned with the development of Back-propagation Neural Network for Bangla Speech Recognition. In this paper, ten bangla digits were recorded from ten speakers and have been recognized. The features of th…

speech-recognitionSpeech Recognition