paper-with-me

홈 › Papers

AFRIDOC-MT: Document-level MT Corpus for African Languages

2025-01-10 · Jesujoba O. Alabi, Israel Abebe Azime, Miaoran Zhang, Cristina España-Bonet, Rachel Bawden, Dawei Zhu, David Ifeoluwa Adelani, Clement Oyeleke Odoje, Idris Akinade, Iffat Maab, Davis David, Shamsuddeen Hassan Muhammad, Neo Putini, David O. Ademuyiwa, Andrew Caines, Dietrich Klakow

This paper introduces AFRIDOC-MT, a document-level multi-parallel translation dataset covering English and five African languages: Amharic, Hausa, Swahili, Yor\`ub\'a, and Zulu. The dataset comprises 334 health and 271 information technology news documents, all human-translated from English to these languages. We conduct document-level translation benchmark experiments by evaluating neural machine translation (NMT) models and large language models (LLMs) for translations between English and these languages, at both the sentence and pseudo-document levels. These outputs are realigned to form complete documents for evaluation. Our results indicate that NLLB-200 achieved the best average performance among the standard NMT models, while GPT-4o outperformed general-purpose LLMs. Fine-tuning selected models led to substantial performance gains, but models trained on sentences struggled to generalize effectively to longer documents. Furthermore, our analysis reveals that some LLMs exhibit issues such as under-generation, repetition of words or phrases, and off-target translations, especially for African languages.

📄 PDF Abstract BibTeX arXiv:2501.06374

Code (1)

masakhane-io/afridoc-mt 공식 구현

Tasks

Machine TranslationNMTSentenceTranslation

Similar Papers 제목 키워드 기반

AfriScience-MT: Towards Decolonizing Science in Africa through Text Translation

2026-05-28 · Idris Abdulmumin, Tajuddeen Gwadabe, Shamsuddeen Hassan Muhammad, David Ifeoluwa Adelani 외 arxiv

The dominance of colonial languages in African education and scientific communication limits how hundreds of millions of speakers of African languages access and produce scientific knowledge. A core obstacle is the lack …

Machine Translation

The African Language Tax: Quantifying the Cost, Latency, and Context Penalty of Tokenizing African Languages in Frontier LLMs

2026-06-23 · Olaoye Anthony Somide arxiv

Commercial large language models bill, scale latency, and budget context per token. Yet tokenizers assign more subword tokens to the same meaning in some languages than in others, so speakers of languages with high token…

Open but Incompatible: A License Compatibility Analysis of Corpora for Low-Resource African Languages

2026-06-27 · Ernst van Gassen arxiv

Creative Commons licenses dominate African NLP corpus releases, but their compatibility rules are rarely applied. CC-BY-SA and CC-BY-NC cannot be combined in a single published dataset; a NoDerivs clause silently prohibi…

Lanfrica: A Participatory Approach to Documenting Machine Translation Research on African Languages

2020-08-03 · Chris C. Emezue, Bonaventure F. P. Dossou

Over the years, there have been campaigns to include the African languages in the growing research on machine translation (MT) in particular, and natural language processing (NLP) in general. Africa has the highest langu…

DiversityMachine TranslationTranslation

Towards a parallel corpus of Portuguese and the Bantu language Emakhuwa of Mozambique

2021-04-12 · Felermino D. M. A. Ali, Andrew Caines, Jaimito L. A. Malavi

Major advancement in the performance of machine translation models has been made possible in part thanks to the availability of large-scale parallel corpora. But for most languages in the world, the existence of such cor…

Machine TranslationSentenceTranslation