paper-with-me

Papers

Multilingual Bottleneck Features for Improving ASR Performance of Code-Switched Speech in Under-Resourced Languages

2020-10-31 · Trideba Padhi, Astik Biswas, Febe De Wet, Ewald van der Westhuizen, Thomas Niesler

In this work, we explore the benefits of using multilingual bottleneck features (mBNF) in acoustic modelling for the automatic speech recognition of code-switched (CS) speech in African languages. The unavailability of annotated corpora in the languages of interest has always been a primary challenge when developing speech recognition systems for this severely under-resourced type of speech. Hence, it is worthwhile to investigate the potential of using speech corpora available for other better-resourced languages to improve speech recognition performance. To achieve this, we train a mBNF extractor using nine Southern Bantu languages that form part of the freely available multilingual NCHLT corpus. We append these mBNFs to the existing MFCCs, pitch features and i-vectors to train acoustic models for automatic speech recognition (ASR) in the target code-switched languages. Our results show that the inclusion of the mBNF features leads to clear performance improvements over a baseline trained without the mBNFs for code-switched English-isiZulu, English-isiXhosa, English-Sesotho and English-Setswana speech.

📄 PDF Abstract BibTeX arXiv:2011.03118

Code (1)

ewaldvdw/kaldi/tree/mbnf_cs2020/egs/nchlt_multi_bnfs/s5 공식 구현

Tasks

Acoustic ModellingAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Codeswitched Sentence Creation using Dependency Parsing

2020-12-05 · Dhruval Jain, Arun D Prabhu, Shubham Vatsal, Gopi Ramena 외

Codeswitching has become one of the most common occurrences across multilingual speakers of the world, especially in countries like India which encompasses around 23 official languages with the number of bilingual speake…

Dependency ParsingSentence

Multilingual Named Entity Recognition on Spanish-English Code-switched Tweets using Support Vector Machines

2018-07-01 · WS 2018 7 · Daniel Claeser, Samantha Kent, Dennis Felske

This paper describes our system submission for the ACL 2018 shared task on named entity recognition (NER) in code-switched Twitter data. Our best result (F1 = 53.65) was obtained using a Support Vector Machine (SVM) with…

Entity LinkingMultilingual Named Entity Recognitionnamed-entity-recognitionNamed Entity Recognition+2

GLUECoS: An Evaluation Benchmark for Code-Switched NLP

2020-07-01 · ACL 2020 6 · Simran Khanuja, D, S apat, ipan 외

Code-switching is the use of more than one language in the same conversation or utterance. Recently, multilingual contextual embedding models, trained on multiple monolingual corpora, have shown promising results on cros…

Language Identificationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+5

GLUECoS : An Evaluation Benchmark for Code-Switched NLP

2020-04-26 · Simran Khanuja, Sandipan Dandapat, Anirudh Srinivasan, Sunayana Sitaram 외

Code-switching is the use of more than one language in the same conversation or utterance. Recently, multilingual contextual embedding models, trained on multiple monolingual corpora, have shown promising results on cros…

Language Identificationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+5

Code-switched inspired losses for generic spoken dialog representations

2021-08-27 · Emile Chapuis, Pierre Colombo, Matthieu Labeau, Chloe Clavel

Spoken dialog systems need to be able to handle both multiple languages and multilinguality inside a conversation (\textit{e.g} in case of code-switching). In this work, we introduce new pretraining losses tailored to le…

Retrieval