Jejueo Datasets for Machine Translation and Speech Synthesis
Jejueo was classified as critically endangered by UNESCO in 2010. Although diverse efforts to revitalize it have been made, there have been few computational approaches. Motivated by this, we construct two new Jejueo datasets: Jejueo Interview Transcripts (JIT) and Jejueo Single Speaker Speech (JSS). The JIT dataset is a parallel corpus containing 170k+ Jejueo-Korean sentences, and the JSS dataset consists of 10k high-quality audio files recorded by a native Jejueo speaker and a transcript file. Subsequently, we build neural systems of machine translation and speech synthesis using them. All resources are publicly available via our GitHub repository. We hope that these datasets will attract interest of both language and machine learning communities.
Code (1)
Tasks
Machine TranslationSpeech SynthesisTranslationSimilar Papers 제목 키워드 기반
Improving Jejueo-Korean Translation With Cross-Lingual Pretraining Using Japanese and Korean
Jejueo is a critically endangered language spoken on Jeju Island and is closely related to but mutually unintelligible with Korean. Parallel data between Jejueo and Korean is scarce, and translation between the two langu…
Machine TranslationTranslationPreparing an Endangered Language for the Digital Age: The Case of Judeo-Spanish
We develop machine translation and speech synthesis systems to complement the efforts of revitalizing Judeo-Spanish, the exiled language of Sephardic Jews, which survived for centuries, but now faces the threat of extinc…
Machine TranslationSpeech Synthesistext-to-speechText to Speech+2A Survey of Voice Translation Methodologies - Acoustic Dialect Decoder
Speech Translation has always been about giving source text or audio input and waiting for system to give translated output in desired form. In this paper, we present the Acoustic Dialect Decoder (ADD) - a voice to voice…
DecoderSentenceSpeech SynthesisSurvey+1Simultaneous Speech-to-Speech Translation System with Neural Incremental ASR, MT, and TTS
This paper presents a newly developed, simultaneous neural speech-to-speech translation system and its evaluation. The system consists of three fully-incremental neural processing modules for automatic speech recognition…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine TranslationSimultaneous Speech-to-Speech Translation+8German-Arabic Speech-to-Speech Translation for Psychiatric Diagnosis
In this paper we present the natural language processing components of our German-Arabic speech-to-speech translation system which is being deployed in the context of interpretation during psychiatric, diagnostic intervi…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderDiagnostic+7