paper-with-me

Papers

Challenging the Boundaries of Speech Recognition: The MALACH Corpus

2019-08-09 · Michael Picheny, Zóltan Tüske, Brian Kingsbury, Kartik Audhkhasi, Xiaodong Cui, George Saon

There has been huge progress in speech recognition over the last several years. Tasks once thought extremely difficult, such as SWITCHBOARD, now approach levels of human performance. The MALACH corpus (LDC catalog LDC2012S05), a 375-Hour subset of a large archive of Holocaust testimonies collected by the Survivors of the Shoah Visual History Foundation, presents significant challenges to the speech community. The collection consists of unconstrained, natural speech filled with disfluencies, heavy accents, age-related coarticulations, un-cued speaker and language switching, and emotional speech - all still open problems for speech recognition systems. Transcription is challenging even for skilled human annotators. This paper proposes that the community place focus on the MALACH corpus to develop speech recognition systems that are more robust with respect to accents, disfluencies and emotional speech. To reduce the barrier for entry, a lexicon and training and testing setups have been created and baseline results using current deep learning technologies are presented. The metadata has just been released by LDC (LDC2019S11). It is hoped that this resource will enable the community to build on top of these baselines so that the extremely important information in these and related oral histories becomes accessible to a wider audience.

📄 PDF Abstract BibTeX arXiv:1908.03455

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Transformer-based Automatic Speech Recognition of Formal and Colloquial Czech in MALACH Project

2022-06-15 · Jan Lehečka, Josef V. Psutka, Josef Psutka

Czech is a very specific language due to its large differences between the formal and the colloquial form of speech. While the formal (written) form is used mainly in official documents, literature, and public speeches, …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Formspeech-recognition+1

Exploring Capabilities of Monolingual Audio Transformers using Large Datasets in Automatic Speech Recognition of Czech

2022-06-15 · Jan Lehečka, Jan Švec, Aleš Pražák, Josef V. Psutka

In this paper, we present our progress in pretraining Czech monolingual audio transformers from a large dataset containing more than 80 thousand hours of unlabeled speech, and subsequently fine-tuning the model on automa…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Long-span language modeling for speech recognition

2019-11-11 · Sarangarajan Parthasarathy, William Gale, Xie Chen, George Polovets 외

We explore neural language modeling for speech recognition where the context spans multiple sentences. Rather than encode history beyond the current sentence using a cache of words or document-level features, we focus ou…

Language ModelingLanguage ModellingRe-RankingSentence+2

Diagonal State Space Augmented Transformers for Speech Recognition

2023-02-27 · George Saon, Ankit Gupta, Xiaodong Cui

We improve on the popular conformer architecture by replacing the depthwise temporal convolutions with diagonal state space (DSS) models. DSS is a recently introduced variant of linear RNNs obtained by discretizing a lin…

speech-recognitionSpeech Recognition

RadioTalk: a large-scale corpus of talk radio transcripts

2019-07-16 · Doug Beeferman, William Brannon, Deb Roy

We introduce RadioTalk, a corpus of speech recognition transcripts sampled from talk radio broadcasts in the United States between October of 2018 and March of 2019. The corpus is intended for use by researchers in the f…

Descriptivespeech-recognitionSpeech Recognition