paper-with-me

Papers

CPLM, a Parallel Corpus for Mexican Languages: Development and Interface

2020-05-01 · LREC 2020 5 · Gerardo Sierra Mart{\'\i}nez, Cynthia Monta{\~n}o, Gemma Bel-Enguix, Diego C{\'o}rdova, Margarita Mota Montoya

Mexico is a Spanish speaking country that has a great language diversity, with 68 linguistic groups and 364 varieties. As they face a lack of representation in education, government, public services and media, they present high levels of endangerment. Due to the lack of data available on social media and the internet, few technologies have been developed for these languages. To analyze different linguistic phenomena in the country, the Language Engineering Group developed the Corpus Paralelo de Lenguas Mexicanas (CPLM) [The Mexican Languages Parallel Corpus], a collaborative parallel corpus for the low-resourced languages of Mexico. The CPLM aligns Spanish with six indigenous languages: Maya, Ch{'}ol, Mazatec, Mixtec, Otomi, and Nahuatl. First, this paper describes the process of building the CPLM: text searching, digitalization and alignment process. Furthermore, we present some difficulties regarding dialectal and orthographic variations. Second, we present the interface and types of searching as well as the use of filters.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Diversity

Similar Papers 제목 키워드 기반

Parallel Corpus for Indigenous Language Translation: Spanish-Mazatec and Spanish-Mixtec

2023-05-27 · Atnafu Lambebo Tonja, Christian Maldonado-Sifuentes, David Alejandro Mendoza Castillo, Olga Kolesnikova 외

In this paper, we present a parallel Spanish-Mazatec and Spanish-Mixtec corpus for machine translation (MT) tasks, where Mazatec and Mixtec are two indigenous Mexican languages. We evaluated the usability of the collecte…

Few-Shot LearningMachine TranslationTransfer LearningTranslation

Automatic Detection of Offensive Language in Social Media: Defining Linguistic Criteria to build a Mexican Spanish Dataset

2020-05-01 · LREC 2020 5 · Mar{\'\i}a Jos{\'e} D{\'\i}az-Torres, Paulina Alej Mor{\'a}n-M{\'e}ndez, ra, Luis Villasenor-Pineda 외

Phenomena such as bullying, homophobia, sexism and racism have transcended to social networks, motivating the development of tools for their automatic detection. The challenge becomes greater for languages rich in popula…

Abusive Language

JW300: A Wide-Coverage Parallel Corpus for Low-Resource Languages

2019-07-01 · ACL 2019 7 · {\v{Z}}eljko Agi{\'c}, Ivan Vuli{\'c}

Viable cross-lingual transfer critically depends on the availability of parallel texts. Shortage of such resources imposes a development and evaluation bottleneck in multilingual processing. We introduce JW300, a paralle…

Cross-Lingual Transfer

The AMARA Corpus: Building Parallel Language Resources for the Educational Domain

2014-05-01 · LREC 2014 5 · Ahmed Abdelali, Francisco Guzman, Hassan Sajjad, Stephan Vogel

This paper presents the AMARA corpus of on-line educational content: a new parallel corpus of educational video subtitles, multilingually aligned for 20 languages, i.e. 20 monolingual corpora and 190 parallel corpora. Th…

Machine TranslationTranslation

A Parallel Corpus for Evaluating Machine Translation between Arabic and European Languages

2017-04-01 · EACL 2017 4 · Nizar Habash, Nasser Zalmout, Dima Taji, Hieu Hoang 외

We present Arab-Acquis, a large publicly available dataset for evaluating machine translation between 22 European languages and Arabic. Arab-Acquis consists of over 12,000 sentences from the JRC-Acquis (Acquis Communauta…

BenchmarkingMachine TranslationTranslation