paper-with-me

Papers

EduMT: Developing Machine Translation System for Educational Content in Indian Languages

2021-12-01 · ICON 2021 12 · Ramakrishna Appicharla, Asif Ekbal, Pushpak Bhattacharyya

In this paper, we explore various approaches to build Hindi to Bengali Neural Machine Translation (NMT) systems for the educational domain. Translation of educational content poses several challenges, such as unavailability of gold standard data for model building, extensive uses of domain-specific terms, as well as the presence of noise in the form of spontaneous speech as the corpus is prepared from subtitle data and noise due to the process of corpus creation through back-translation. We create an educational parallel corpus by crawling lecture subtitles and translating them into Hindi and Bengali using Google translate. We also create a clean parallel corpus by post-editing synthetic corpus via annotation and crowd-sourcing. We build NMT systems on the prepared corpus with domain adaptation objectives. We also explore data augmentation methods by automatically cleaning synthetic corpus and using it to further train the models. We experiment with combining domain adaptation objective with multilingual NMT. We report BLEU and TER scores of all the models on a manually created Hindi-Bengali educational testset. Our experiments show that the multilingual domain adaptation model outperforms all the other models by achieving 34.8 BLEU and 0.466 TER scores.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationDomain AdaptationMachine TranslationNMTTranslation

Similar Papers 제목 키워드 기반

Mitigating Language Barriers in Education: Developing Multilingual Digital Learning Materials with Machine Translation

2025-09-11 · Lucie Poláková, Martin Popel, Věra Kloudová, Michal Novák 외 arxiv

The EdUKate project combines digital education, linguistics, translation studies, and machine translation to develop multilingual learning materials for Czech primary and secondary schools. Launched through collaboration…

Machine Translation

The AMARA Corpus: Building Parallel Language Resources for the Educational Domain

2014-05-01 · LREC 2014 5 · Ahmed Abdelali, Francisco Guzman, Hassan Sajjad, Stephan Vogel

This paper presents the AMARA corpus of on-line educational content: a new parallel corpus of educational video subtitles, multilingually aligned for 20 languages, i.e. 20 monolingual corpora and 190 parallel corpora. Th…

Machine TranslationTranslation

PEACH: A sentence-aligned Parallel English-Arabic Corpus for Healthcare

2025-08-07 · Rania Al-Sabbagh arxiv

This paper introduces PEACH, a sentence-aligned parallel English-Arabic corpus of healthcare texts encompassing patient information leaflets and educational materials. The corpus contains 51,671 parallel sentences, total…

Machine Translation

Applying Automated Machine Translation to Educational Video Courses

2023-01-09 · Linden Wang

We studied the capability of automated machine translation in the online video education space by automatically translating Khan Academy videos with state-of-the-art translation models and applying text-to-speech synthes…

Machine TranslationSpeech Synthesistext-to-speechText to Speech+3

QE Viewer: an Open-Source Tool for Visualization of Machine Translation Quality Estimation Results

2020-11-01 · EAMT 2020 11 · Felipe Soares, Anna Zaretskaya, Diego Bartolome

QE Viewer is a web-based tool for visualizing results of a Machine Translation Quality Estimation (QE) system. It allows users to see information on the predicted post-editing distance (PED) for a given file or sentence,…

Machine TranslationSentenceTranslation