paper-with-me

홈 › Papers

Zero-Shot Morphological Discovery in Low-Resource Bantu Languages via Cross-Lingual Transfer and Unsupervised Clustering

2026-04-24 · Hillary Mutisya, John Mugane arxiv

We present a method for discovering morphological features in low-resource Bantu languages by combining cross-lingual transfer learning with unsupervised clustering. Applied to Giriama (nyf), a language with only 91 labeled paradigms, our pipeline discovers noun class assignments for 2,455 words and identifies two previously undocumented morphological patterns: an a- prefix variant for Class 2 (vowel coalescence - the merger of two adjacent vowels - of wa-, 95.1% consistency) and a contracted k'- prefix (98.5% consistency). External validation on 444 known Giriama verb paradigms confirms 78.2% lemmatization accuracy, while a v3 corpus expansion to 19,624 words (9,014 unique lemmas) achieves 97.3% segmentation and 86.7% lemmatization rates across all major word classes. Our ensemble of transfer learning from Swahili and unsupervised clustering, combined via weighted voting, exploits complementary strengths: transfer excels at cognate detection (leveraging ~60% vocabulary overlap) while clustering discovers language-specific innovations invisible to transfer. We release all code and discovered lexicons to support morphological documentation for low-resource Bantu languages.

📄 PDF Abstract BibTeX arXiv:2604.22723

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Lingual TransferTransfer Learning

Similar Papers 제목 키워드 기반

Shona spaCy: A Morphological Analyzer for an Under-Resourced Bantu Language

2025-11-12 · Happymore Masoka arxiv

Despite rapid advances in multilingual natural language processing (NLP), the Bantu language Shona remains under-served in terms of morphological analysis and language-aware tools. This paper presents Shona spaCy, an ope…

Neural Recovery of Historical Lexical Structure in Bantu Languages from Modern Data

2026-04-24 · Hillary Mutisya, John Mugane arxiv

We investigate whether neural models trained exclusively on modern morphological data can recover cross-lingual lexical structure consistent with historical reconstruction. Using BantuMorph v7, a transformer over Bantu m…

Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition

2026-06-30 · Kesego Mokgosi, Vukosi Marivate, Sitwala Mundia, Unarine Netshifhefhe 외 arxiv

Southern Bantu languages are spoken by over 80 million people, yet current foundation ASR models still produce zero-shot WER above 100%, which limits practical use in education and public services. We addressed this gap …

Speech Recognition

A Very Low Resource Language Speech Corpus for Computational Language Documentation Experiments

2017-10-10 · LREC 2018 5 · P. Godard, G. Adda, M. Adda-Decker, J. Benjumea 외

Most speech and language technologies are trained with massive amounts of speech and text information. However, most of the world languages do not have such resources or stable orthography. Systems constructed under thes…

Low-Resource Neural Machine Translation for Southern African Languages

2021-04-01 · Evander Nyoni, Bruce A. Bassett

Low-resource African languages have not fully benefited from the progress in neural machine translation because of a lack of data. Motivated by this challenge we compare zero-shot learning, transfer learning and multilin…

Low Resource Neural Machine TranslationLow-Resource Neural Machine TranslationMachine TranslationSentence+3