paper-with-me

홈 › Papers

Mafoko: Structuring and Building Open Multilingual Terminologies for South African NLP

2025-08-05 · Vukosi Marivate, Isheanesu Dzingirai, Fiskani Banda, Richard Lastrucci, Thapelo Sindane, Keabetswe Madumo, Kayode Olaleye, Abiodun Modupe, Unarine Netshifhefhe, Herkulaas Combrink, Mohlatlego Nakeng, Matome Ledwaba arxiv

The critical lack of structured terminological data for South Africa's official languages hampers progress in multilingual NLP, despite the existence of numerous government and academic terminology lists. These valuable assets remain fragmented and locked in non-machine-readable formats, rendering them unusable for computational research and development. Mafoko addresses this challenge by systematically aggregating, cleaning, and standardising these scattered resources into open, interoperable datasets. We introduce the foundational Mafoko dataset, released under the equitable, Africa-centered NOODL framework. To demonstrate its immediate utility, we integrate the terminology into a Retrieval-Augmented Generation (RAG) pipeline. Experiments show substantial improvements in the accuracy and domain-specific consistency of English-to-Tshivenda machine translation for large language models. Mafoko provides a scalable foundation for developing robust and equitable NLP technologies, ensuring South Africa's rich linguistic diversity is represented in the digital age.

📄 PDF Abstract BibTeX arXiv:2508.03529

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Translation

Similar Papers 제목 키워드 기반

Multilingualization of Medical Terminology: Semantic and Structural Embedding Approaches

2020-05-01 · LREC 2020 5 · Long-Huei Chen, Kyo Kageura

The multilingualization of terminology is an essential step in the translation pipeline, to ensure the correct transfer of domain-specific concepts. Many institutions and language service providers construct and maintain…

Term ExtractionTranslation

GENIE: Generative Note Information Extraction model for structuring EHR data

2025-01-30 · Huaiyuan Ying, Hongyi Yuan, Jinsen Lu, Zitian Qu 외

Electronic Health Records (EHRs) hold immense potential for advancing healthcare, offering rich, longitudinal data that combines structured information with valuable insights from unstructured clinical notes. However, th…

AttributeAttribute Extraction

LIDIOMS: A Multilingual Linked Idioms Data Set

2018-02-22 · LREC 2018 5 · Diego Moussallem, Mohamed Ahmed Sherif, Diego Esteves, Marcos Zampieri 외

In this paper, we describe the LIDIOMS data set, a multilingual RDF representation of idioms currently containing five languages: English, German, Italian, Portuguese, and Russian. The data set is intended to support nat…

Language Technologies for the Creation of Multilingual Terminologies. Lessons Learned from the SSHOC Project

2022-06-01 · LREC 2022 6 · Federica Gamba, Francesca Frontini, Daan Broeder, Monica Monachini

This paper is framed in the context of the SSHOC project and aims at exploring how Language Technologies can help in promoting and facilitating multilingualism in the Social Sciences and Humanities (SSH). Although most S…

Machine TranslationTranslationvalid

From Linguistic Resources to Ontology-Aware Terminologies: Minding the Representation Gap

2020-05-01 · LREC 2020 5 · Giulia Speranza, Maria Pia di Buono, Johanna Monti, Federico Sangati

Terminological resources have proven crucial in many applications ranging from Computer-Aided Translation tools to authoring softwares and multilingual and cross-lingual information retrieval systems. Nonetheless, with t…

Cross-Lingual Information RetrievalInformation RetrievalRetrievalTranslation