paper-with-me

홈 › Papers

$π$-yalli: un nouveau corpus pour le nahuatl

2024-12-20 · Juan-Manuel Torres-Moreno, Juan-José Guzmán-Landa, Graham Ranger, Martha Lorena Avendaño Garrido, Miguel Figueroa-Saavedra, Ligia Quintana-Torres, Carlos-Emiliano González-Gallardo, Elvys Linhares Pontes, Patricia Velázquez Morales, Luis-Gil Moreno Jiménez

The NAHU$^2$ project is a Franco-Mexican collaboration aimed at building the $\pi$-YALLI corpus adapted to machine learning, which will subsequently be used to develop computer resources for the Nahuatl language. Nahuatl is a language with few computational resources, even though it is a living language spoken by around 2 million people. We have decided to build $\pi$-YALLI, a corpus that will enable to carry out research on Nahuatl in order to develop Language Models (LM), whether dynamic or not, which will make it possible to in turn enable the development of Natural Language Processing (NLP) tools such as: a) a grapheme unifier, b) a word segmenter, c) a POS grammatical analyser, d) a content-based Automatic Text Summarization; and possibly, e) a translator translator (probabilistic or learning-based).

📄 PDF Abstract BibTeX arXiv:2412.15821

Code (0)

등록된 구현이 없습니다.

Tasks

POSText Summarization

Similar Papers 제목 키워드 기반

Cr\'eation d'un nouveau treebank \`a partir de quatri\`emes de couverture

2015-06-01 · JEPTALNRECITAL 2015 6 · Philippe Blache, Gr{\'e}goire Moncheuil, St{\'e}phane Rauzy, Marie-Laure Gu{\'e}not

Nous pr{\'e}sentons ici 4-couv, un nouveau corpus arbor{\'e} d{'}environ 3 500 phrases, constitu{\'e} d{'}un ensemble de quatri{\`e}mes de couverture, {\'e}tiquet{\'e} et analys{\'e} automatiquement puis corrig{\'e} et v…

valid

A symbolic Perl algorithm for the unification of Nahuatl word spellings

2025-11-24 · Juan-José Guzmán-Landa, Jesús Vázquez-Osorio, Juan-Manuel Torres-Moreno, Ligia Quintana Torres 외 arxiv

In this paper, we describe a symbolic model for the automatic orthographic unification of Nawatl text documents. Our model is based on algorithms that we have previously used to analyze sentences in Nawatl, and on the co…

A First Context-Free Grammar Applied to Nawatl Corpora Augmentation

2025-10-06 · Juan-José Guzmán-Landa, Juan-Manuel Torres-Moreno, Miguel Figueroa-Saavedra, Ligia Quintana-Torres 외 arxiv

In this article we introduce a context-free grammar (CFG) for the Nawatl language. Nawatl (or Nahuatl) is an Amerindian language of the $π$-language type, i.e. a language with few digital resources, in which the corpora …

Un corpus d'\'evaluation pour un syst\`eme de simplification discursive (An Evaluation Corpus for Automatic Discourse Simplification)

2020-06-01 · JEPTALNRECITAL 2020 6 · Rodrigo Wilkens, Amalia Todirascu

Nous pr{\'e}sentons un nouveau corpus simplifi{\'e}, disponible en fran{\c{c}}ais pour l{'}{\'e}valuation d{'}un syst{\`e}me de simplification discursive. Ce syst{\`e}me utilise des cha{\^\i}nes de r{\'e}f{\'e}rence pour…

ACCOLÉ : Annotation Collaborative d’erreurs de traduction pour COrpus aLignÉs, multi-cibles, et Annotation d’Expressions Poly-lexicales (ACCOLÉ: A Collaborative Platform of Error Annotation for Aligned)

2021-06-01 · JEP/TALN/RECITAL 2021 6 · Emmanuelle Esperança-Rodier, Francis Brunet-Manquat

Cette démonstration présente les avancées d’ACCOLÉ (Annotation Collaborative d’erreurs de traduction pour COrpus aLignÉs), qui en plus de proposer une gestion simplifiée des corpus et des typologies d’erreurs, l’annotati…