paper-with-me

Papers

Building a Functional Machine Translation Corpus for Kpelle

2025-05-24 · Kweku Andoh Yamoah, Jackson Weako, Emmanuel J. Dorley

In this paper, we introduce the first publicly available English-Kpelle dataset for machine translation, comprising over 2000 sentence pairs drawn from everyday communication, religious texts, and educational materials. By fine-tuning Meta's No Language Left Behind(NLLB) model on two versions of the dataset, we achieved BLEU scores of up to 30 in the Kpelle-to-English direction, demonstrating the benefits of data augmentation. Our findings align with NLLB-200 benchmarks on other African languages, underscoring Kpelle's potential for competitive performance despite its low-resource status. Beyond machine translation, this dataset enables broader NLP tasks, including speech recognition and language modelling. We conclude with a roadmap for future dataset expansion, emphasizing orthographic consistency, community-driven validation, and interdisciplinary collaboration to advance inclusive language technology development for Kpelle and other low-resourced Mande languages.

📄 PDF Abstract BibTeX arXiv:2505.18905

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationLanguage ModellingMachine TranslationSentencespeech-recognitionSpeech RecognitionTranslation

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Towards Santali Linguistic Inclusion: Building the First Santali-to-English Translation Model using mT5 Transformer and Data Augmentation

2024-11-29 · Syed Mohammed Mostaque Billah, Ateya Ahmed Subarna, Sudipta Nandi Sarna, Ahmad Shawkat Wasit 외

Around seven million individuals in India, Bangladesh, Bhutan, and Nepal speak Santali, positioning it as nearly the third most commonly used Austroasiatic language. Despite its prominence among the Austroasiatic languag…

Data AugmentationMachine TranslationTransfer LearningTranslation

Parallel resources for Tunisian Arabic Dialect Translation

2020-12-01 · COLING (WANLP) 2020 12 · Saméh Kchaou, Rahma Boujelbane, Lamia Hadrich-Belguith

The difficulty of processing dialects is clearly observed in the high cost of building representative corpus, in particular for machine translation. Indeed, all machine translation systems require a huge amount and good …

Data AugmentationMachine TranslationManagementSentence+1

KC4MT: A High-Quality Corpus for Multilingual Machine Translation

2022-06-01 · LREC 2022 6 · Vinh Van Nguyen, Ha Nguyen, Huong Thanh Le, Thai Phuong Nguyen 외

The multilingual parallel corpus is an important resource for many applications of natural language processing (NLP). For machine translation, the size and quality of the training corpus mainly affects the quality of the…

Machine TranslationSentenceTranslationVocal Bursts Intensity Prediction

The AMARA Corpus: Building Parallel Language Resources for the Educational Domain

2014-05-01 · LREC 2014 5 · Ahmed Abdelali, Francisco Guzman, Hassan Sajjad, Stephan Vogel

This paper presents the AMARA corpus of on-line educational content: a new parallel corpus of educational video subtitles, multilingually aligned for 20 languages, i.e. 20 monolingual corpora and 190 parallel corpora. Th…

Machine TranslationTranslation

Ukrainian-to-English folktale corpus: Parallel corpus creation and augmentation for machine translation in low-resource languages

2024-10-14 · AMTA 2022 9 · Olena Burda-Lassen

Folktales are linguistically very rich and culturally significant in understanding the source language. Historically, only human translation has been used for translating folklore. Therefore, the number of translated tex…

Machine TranslationSentenceTranslation