paper-with-me

Papers

A Multi Purpose and Large Scale Speech Corpus in Persian and English for Speaker and Speech Recognition: the DeepMine Database

2019-12-08 · Hossein Zeinali, Lukáš Burget, Jan "Honza'' Černocký

DeepMine is a speech database in Persian and English designed to build and evaluate text-dependent, text-prompted, and text-independent speaker verification, as well as Persian speech recognition systems. It contains more than 1850 speakers and 540 thousand recordings overall, more than 480 hours of speech are transcribed. It is the first public large-scale speaker verification database in Persian, the largest public text-dependent and text-prompted speaker verification database in English, and the largest public evaluation dataset for text-independent speaker verification. It has a good coverage of age, gender, and accents. We provide several evaluation protocols for each part of the database to allow for research on different aspects of speaker verification. We also provide the results of several experiments that can be considered as baselines: HMM-based i-vectors for text-dependent speaker verification, and HMM-based as well as state-of-the-art deep neural network based ASR. We demonstrate that the database can serve for training robust ASR models.

📄 PDF Abstract BibTeX arXiv:1912.03627

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Verificationspeech-recognitionSpeech RecognitionText-Dependent Speaker VerificationText-Independent Speaker Verification

Similar Papers 제목 키워드 기반

A Multi-Purpose Audio-Visual Corpus for Multi-Modal Persian Speech Recognition: the Arman-AV Dataset

2023-01-21 · Javad Peymanfard, Samin Heydarian, Ali Lashini, Hossein Zeinali 외

In recent years, significant progress has been made in automatic lip reading. But these methods require large-scale datasets that do not exist for many low-resource languages. In this paper, we have presented a new multi…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Lip Reading+4

Common Voice: A Massively-Multilingual Speech Corpus

2019-12-13 · LREC 2020 5 · Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty 외

The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Identificationspeech-recognition+3

Polish Read Speech Corpus for Speech Tools and Services

2017-06-01 · Danijel Koržinek, Krzysztof Marasek, Łukasz Brocki, Krzysztof Wołk

This paper describes the speech processing activities conducted at the Polish consortium of the CLARIN project. The purpose of this segment of the project was to develop specific tools that would allow for automatic and …

Action DetectionActivity DetectionGrapheme-to-Phoneme ConversionKeyword Spotting+3

A Game with a Purpose for Automatic Detection of Children's Speech Disabilities using Limited Speech Resources

2017-09-01 · RANLP 2017 9 · Reem Salem, Mohamed Elmahdy, Slim Abdennadher, Injy Hamed

Speech therapists and researchers are becoming more concerned with the use of computer-based systems in the therapy of speech disorders. In this paper, we propose a computer-based game with a purpose (GWAP) for speech th…

Information Retrieval

JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis

2017-10-28 · Ryosuke Sonobe, Shinnosuke Takamichi, Hiroshi Saruwatari

Thanks to improvements in machine learning techniques including deep learning, a free large-scale speech corpus that can be shared between academic institutions and commercial companies has an important role. However, su…

BIG-bench Machine LearningSpeech Synthesis