paper-with-me

Papers

Crowd-sourced Phrase-Based Tokenization for Low-Resourced Neural Machine Translation: The case of Fon Language

2021-01-01 · Bonaventure F. P. Dossou, Chris Chinenye Emezue

Building effective neural machine translation (NMT) models for very low-resourced and morphologically rich African indigenous languages is an open challenge. Besides the issue of finding available resources for them, a lot of work is put into preprocessing and tokenization. Recent studies have shown that standard tokenization methods do not always adequately deal with the grammatical, diacritical, and tonal properties of some African languages. That, coupled with the extremely low availability of training samples, hinders the production of reliable NMT models. In this paper, using Fon language as a case study, we revisit standard tokenization methods and introduce Word-Expressions-Based (WEB) tokenization, a human-involved super-words tokenization strategy to create a better representative vocabulary for training.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationNMTTranslation

Similar Papers 제목 키워드 기반

Crowdsourced Phrase-Based Tokenization for Low-Resourced Neural Machine Translation: The Case of Fon Language

2021-03-14 · Bonaventure F. P. Dossou, Chris C. Emezue

Building effective neural machine translation (NMT) models for very low-resourced and morphologically rich African indigenous languages is an open challenge. Besides the issue of finding available resources for them, a l…

Machine TranslationNMTTranslation

Linguistically Informed Tokenization Improves ASR for Underresourced Languages

2025-10-07 · Massimo Daul, Alessio Tosolini, Claire Bowern arxiv

Automatic speech recognition (ASR) is a crucial tool for linguists aiming to perform a variety of language documentation tasks. However, modern ASR systems use data-hungry transformer architectures, rendering them genera…

Speech Recognition

Automatic Bilingual Phrase Dictionary Construction from GIZA++ Output

2022-06-01 · LREC (MWE) 2022 6 · Albina Khusainova, Vitaly Romanov, Adil Khan

Modern encoder-decoder based neural machine translation (NMT) models are normally trained on parallel sentences. Hence, they give best results when translating full sentences rather than sentence parts. Thereby, the task…

DecoderMachine TranslationNMTSentence+1

Deep Learning Transformer Architecture for Named Entity Recognition on Low Resourced Languages: State of the art results

2021-11-01 · Ridewaan Hanslo

This paper reports on the evaluation of Deep Learning (DL) transformer architecture models for Named-Entity Recognition (NER) on ten low-resourced South African (SA) languages. In addition, these DL transformer models we…

BIG-bench Machine LearningChunkingMachine Translationnamed-entity-recognition+5

Toward a Lightweight Solution for Less-resourced Languages: Creating a POS Tagger for Alsatian Using Voluntary Crowdsourcing

2018-05-01 · LREC 2018 5 · Alice Millour, Kar{\"e}n Fort
POS