paper-with-me

Papers

Axolotl: a Web Accessible Parallel Corpus for Spanish-Nahuatl

2016-05-01 · LREC 2016 5 · Ximena Gutierrez-Vasques, Gerardo Sierra, Isaac Hern Pompa, ez

This paper describes the project called Axolotl which comprises a Spanish-Nahuatl parallel corpus and its search interface. Spanish and Nahuatl are distant languages spoken in the same country. Due to the scarcity of digital resources, we describe the several problems that arose when compiling this corpus: most of our sources were non-digital books, we faced errors when digitizing the sources and there were difficulties in the sentence alignment process, just to mention some. The documents of the parallel corpus are not homogeneous, they were extracted from different sources, there is dialectal, diachronical, and orthographical variation. Additionally, we present a web search interface that allows to make queries through the whole parallel corpus, the system is capable to retrieve the parallel fragments that contain a word or phrase searched by a user in any of the languages. To our knowledge, this is the first Spanish-Nahuatl public available digital parallel corpus. We think that this resource can be useful to develop language technologies and linguistic studies for this language pair.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Similar Papers 제목 키워드 기반

Comparing morphological complexity of Spanish, Otomi and Nahuatl

2018-08-13 · WS 2018 8 · Ximena Gutierrez-Vasques, Victor Mijangos

We use two small parallel corpora for comparing the morphological complexity of Spanish, Otomi and Nahuatl. These are languages that belong to different linguistic families, the latter are low-resourced. We take into acc…

BPE vs. Morphological Segmentation: A Case Study on Machine Translation of Four Polysynthetic Languages

2022-03-16 · Findings (ACL) 2022 5 · Manuel Mager, Arturo Oncevay, Elisabeth Mager, Katharina Kann 외

Morphologically-rich polysynthetic languages present a challenge for NLP systems due to data sparsity, and a common strategy to handle this issue is to apply subword segmentation. We investigate a wide variety of supervi…

Machine TranslationSegmentationTranslation

CPLM, a Parallel Corpus for Mexican Languages: Development and Interface

2020-05-01 · LREC 2020 5 · Gerardo Sierra Mart{\'\i}nez, Cynthia Monta{\~n}o, Gemma Bel-Enguix, Diego C{\'o}rdova 외

Mexico is a Spanish speaking country that has a great language diversity, with 68 linguistic groups and 364 varieties. As they face a lack of representation in education, government, public services and media, they prese…

Diversity

Low-resource bilingual lexicon extraction using graph based word embeddings

2017-10-06 · Ximena Gutierrez-Vasques, Victor Mijangos

In this work we focus on the task of automatically extracting bilingual lexicon for the language pair Spanish-Nahuatl. This is a low-resource setting where only a small amount of parallel corpus is available. Most of the…

TranslationWord AlignmentWord Embeddings

$π$-yalli: un nouveau corpus pour le nahuatl

2024-12-20 · Juan-Manuel Torres-Moreno, Juan-José Guzmán-Landa, Graham Ranger, Martha Lorena Avendaño Garrido 외

The NAHU$^2$ project is a Franco-Mexican collaboration aimed at building the $\pi$-YALLI corpus adapted to machine learning, which will subsequently be used to develop computer resources for the Nahuatl language. Nahuatl…

POSText Summarization