paper-with-me

Papers

The AMARA Corpus: Building Parallel Language Resources for the Educational Domain

2014-05-01 · LREC 2014 5 · Ahmed Abdelali, Francisco Guzman, Hassan Sajjad, Stephan Vogel

This paper presents the AMARA corpus of on-line educational content: a new parallel corpus of educational video subtitles, multilingually aligned for 20 languages, i.e. 20 monolingual corpora and 190 parallel corpora. This corpus includes both resource-rich languages such as English and Arabic, and resource-poor languages such as Hindi and Thai. In this paper, we describe the gathering, validation, and preprocessing of a large collection of parallel, community-generated subtitles. Furthermore, we describe the methodology used to prepare the data for Machine Translation tasks. Additionally, we provide a document-level, jointly aligned development and test sets for 14 language pairs, designed for tuning and testing Machine Translation systems. We provide baseline results for these tasks, and highlight some of the challenges we face when building machine translation systems for educational content.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Quality versus Quantity: Building Catalan-English MT Resources

2022-06-01 · SIGUL (LREC) 2022 6 · Ona de Gibert Bonet, Ksenia Kharitonova, Blanca Calvo Figueras, Jordi Armengol-Estapé 외

In this work, we make the case of quality over quantity when training a MT system for a medium-to-low-resource language pair, namely Catalan-English. We compile our training corpus out of existing resources of varying qu…

Cross-Lingual TransferTransfer LearningTranslation

Building a Bilingual Vietnamese-French Named Entity Annotated Corpus through Cross-Linguistic Projection

2015-06-01 · JEPTALNRECITAL 2015 6 · Ngoc Tan Le, Fatiha Sadat

The creation of high-quality named entity annotated resources is time-consuming and an expensive process. Most of the gold standard corpora are available for English but not for less-resourced languages such as Vietnames…

A fully automated and scalable Parallel Data Augmentation for Low Resource Languages using Image and Text Analytics

2025-10-15 · Prawaal Sharma, Navneet Goyal, Poonam Goyal, Vishnupriyan R arxiv

Linguistic diversity across the world creates a disparity with the availability of good quality digital language resources thereby restricting the technological benefits to majority of human population. The lack or absen…

Machine TranslationData Augmentation

A Web Tool for Building Parallel Corpora of Spoken and Sign Languages

2016-05-01 · LREC 2016 5 · Alex Becker, Fabio Kepler, C, Sara eias

In this paper we describe our work in building an online tool for manually annotating texts in any spoken language with SignWriting in any sign language. The existence of such tool will allow the creation of parallel cor…

Translation

Preparing an Endangered Language for the Digital Age: The Case of Judeo-Spanish

2022-05-31 · EURALI (LREC) 2022 6 · Alp Öktem, Rodolfo Zevallos, Yasmin Moslem, Güneş Öztürk 외

We develop machine translation and speech synthesis systems to complement the efforts of revitalizing Judeo-Spanish, the exiled language of Sephardic Jews, which survived for centuries, but now faces the threat of extinc…

Machine TranslationSpeech Synthesistext-to-speechText to Speech+2