paper-with-me

홈 › Papers

A Parallel Corpus Mixtec-Spanish

2019-08-01 · WS 2019 8 · Cynthia Monta{\~n}o, Gerardo Sierra Mart{\'\i}nez, Gemma Bel-Enguix, Helena Gomez

This work is about the compilation process of parallel documents Spanish-Mixtec. There are not many Spanish-Mixec parallel texts and most of the sources are non-digital books. Due to this, we need to face the errors when digitizing the sources and difficulties in sentence alignment, as well as the fact that does not exist a standard orthography. Our parallel corpus consists of sixty texts coming from books and digital repositories. These documents belong to different domains: history, traditional stories, didactic material, recipes, ethnographical de- scriptions of each town and instruction manuals for disease prevention. We have classified this material in five major categories: didactic (6 texts), educative (6 texts), interpretative (7 texts), narrative (39 texts), and poetic (2 texts). The final total of tokens is 49,814 Spanish words and 47,774 Mixtec words. The texts belong to the states of Oaxaca (48 texts), Guerrero (9 texts) and Puebla (3 texts). According to this data, we see that the corpus is unbalanced in what refers to the representation of the different territories. While 55{\%} of speakers are in Oaxaca, 80{\%} of texts come from this region. Guerrero has the 30{\%} of speakers and the 15{\%} of texts and Puebla, with the 15{\%} of the speakers has a representation of the 5{\%} in the corpus.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Similar Papers 제목 키워드 기반

Parallel Corpus for Indigenous Language Translation: Spanish-Mazatec and Spanish-Mixtec

2023-05-27 · Atnafu Lambebo Tonja, Christian Maldonado-Sifuentes, David Alejandro Mendoza Castillo, Olga Kolesnikova 외

In this paper, we present a parallel Spanish-Mazatec and Spanish-Mixtec corpus for machine translation (MT) tasks, where Mazatec and Mixtec are two indigenous Mexican languages. We evaluated the usability of the collecte…

Few-Shot LearningMachine TranslationTransfer LearningTranslation

CPLM, a Parallel Corpus for Mexican Languages: Development and Interface

2020-05-01 · LREC 2020 5 · Gerardo Sierra Mart{\'\i}nez, Cynthia Monta{\~n}o, Gemma Bel-Enguix, Diego C{\'o}rdova 외

Mexico is a Spanish speaking country that has a great language diversity, with 68 linguistic groups and 364 varieties. As they face a lack of representation in education, government, public services and media, they prese…

Diversity

Development of a Guarani - Spanish Parallel Corpus

2020-05-01 · LREC 2020 5 · Luis Chiruzzo, Pedro Amarilla, Adolfo R{\'\i}os, Gustavo Gim{\'e}nez Lugo

This paper presents the development of a Guarani - Spanish parallel corpus with sentence-level alignment. The Guarani sentences of the corpus use the Jopara Guarani dialect, the dialect of Guarani spoken in Paraguay, whi…

Sentence

Axolotl: a Web Accessible Parallel Corpus for Spanish-Nahuatl

2016-05-01 · LREC 2016 5 · Ximena Gutierrez-Vasques, Gerardo Sierra, Isaac Hern Pompa, ez

This paper describes the project called Axolotl which comprises a Spanish-Nahuatl parallel corpus and its search interface. Spanish and Nahuatl are distant languages spoken in the same country. Due to the scarcity of dig…

Sentence

Design and compilation of a specialized Spanish-German parallel corpus

2012-05-01 · LREC 2012 5 · Carla Parra Escart{\'\i}n

This paper discusses the design and compilation of the TRIS corpus, a specialized parallel corpus of Spanish and German texts. It will be used for phraseological research aimed at improving statistical machine translatio…

Machine TranslationSentenceTranslation