paper-with-me

홈 › Papers

Synthetic Data for Neural Machine Translation of Spoken-Dialects

2017-07-01 · IWSLT 2017 12 · Hany Hassan, Mostafa ElAraby, Ahmed Tawfik

In this paper, we introduce a novel approach to generate synthetic data for training Neural Machine Translation systems. The proposed approach transforms a given parallel corpus between a written language and a target language to a parallel corpus between a spoken dialect variant and the target language. Our approach is language independent and can be used to generate data for any variant of the source language such as slang or spoken dialect or even for a different language that is closely related to the source language. The proposed approach is based on local embedding projection of distributed representations which utilizes monolingual embeddings to transform parallel data across language variants. We report experimental results on Levantine to English translation using Neural Machine Translation. We show that the generated data can improve a very large scale system by more than 2.8 Bleu points using synthetic spoken data which shows that it can be used to provide a reliable translation system for a spoken dialect that does not have sufficient parallel data.

📄 PDF Abstract BibTeX arXiv:1707.00079

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Kurdish Interdialect Machine Translation

2017-04-01 · WS 2017 4 · Hossein Hassani

This research suggests a method for machine translation among two Kurdish dialects. We chose the two widely spoken dialects, Kurmanji and Sorani, which are considered to be mutually unintelligible. Also, despite being sp…

Machine TranslationTranslation

Machine Translation of Low-Resource Spoken Dialects: Strategies for Normalizing Swiss German

2017-10-30 · LREC 2018 5 · Pierre-Edouard Honnet, Andrei Popescu-Belis, Claudiu Musat, Michael Baeriswyl

The goal of this work is to design a machine translation (MT) system for a low-resource family of dialects, collectively known as Swiss German, which are widely spoken in Switzerland but seldom written. We collected a si…

DiversityMachine TranslationText NormalizationTranslation

BanglaDialecto: An End-to-End AI-Powered Regional Speech Standardization

2024-11-16 · Md. Nazmus Sadat Samin, Jawad Ibn Ahad, Tanjila Ahmed Medha, Fuad Rahman 외

This study focuses on recognizing Bangladeshi dialects and converting diverse Bengali accents into standardized formal Bengali speech. Dialects, often referred to as regional languages, are distinctive variations of a la…

Machine Translationspeech-recognitionSpeech Recognition

Voices Unheard: NLP Resources and Models for Yorùbá Regional Dialects

2024-06-27 · Orevaoghene Ahia, Anuoluwapo Aremu, Diana Abagyan, Hila Gonen 외

Yor\`ub\'a an African language with roughly 47 million speakers encompasses a continuum with several dialects. Recent efforts to develop NLP technologies for African languages have focused on their standard dialects, res…

Automatic Speech RecognitionMachine Translationspeech-recognitionSpeech Recognition+3

The SADID Evaluation Datasets for Low-Resource Spoken Language Machine Translation of Arabic Dialects

2020-12-01 · COLING 2020 8 · Wael Abid

Low-resource Machine Translation recently gained a lot of popularity, and for certain languages, it has made great strides. However, it is still difficult to track progress in other languages for which there is no public…

Machine TranslationTranslation