Improving Code-switching Language Modeling with Artificially Generated Texts using Cycle-consistent Adversarial Networks
This paper presents our latest effort on improving Code-switching language models that suffer from data scarcity. We investigate methods to augment Code-switching training text data by artificially generating them. Concretely, we propose a cycle-consistent adversarial networks based framework to transfer monolingual text into Code-switching text, considering Code-switching as a speaking style. Our experimental results on the SEAME corpus show that utilising artificially generated Code-switching text data improves consistently the language model as well as the automatic speech recognition performance.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modellingspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Arabic Code-Switching Speech Recognition using Monolingual Data
Code-switching in automatic speech recognition (ASR) is an important challenge due to globalization. Recent research in multilingual ASR shows potential improvement over monolingual systems. We study key issues related t…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionTraining a code-switching language model with monolingual data
A lack of code-switching data complicates the training of code-switching (CS) language models. We propose an approach to train such CS language models on monolingual data only. By constraining and normalizing the output …
Language ModelingLanguage ModellingTranslationWord TranslationGLUECoS: An Evaluation Benchmark for Code-Switched NLP
Code-switching is the use of more than one language in the same conversation or utterance. Recently, multilingual contextual embedding models, trained on multiple monolingual corpora, have shown promising results on cros…
Language Identificationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+5GLUECoS : An Evaluation Benchmark for Code-Switched NLP
Code-switching is the use of more than one language in the same conversation or utterance. Recently, multilingual contextual embedding models, trained on multiple monolingual corpora, have shown promising results on cros…
Language Identificationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+5Integrating Knowledge in End-to-End Automatic Speech Recognition for Mandarin-English Code-Switching
Code-Switching (CS) is a common linguistic phenomenon in multilingual communities that consists of switching between languages while speaking. This paper presents our investigations on end-to-end speech recognition for M…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language IdentificationLanguage Modeling+3