paper-with-me

홈 › Papers

Designing and Contextualising Probes for African Languages

2025-05-15 · Wisdom Aduah, Francois Meyer

Pretrained language models (PLMs) for African languages are continually improving, but the reasons behind these advances remain unclear. This paper presents the first systematic investigation into probing PLMs for linguistic knowledge about African languages. We train layer-wise probes for six typologically diverse African languages to analyse how linguistic features are distributed. We also design control tasks, a way to interpret probe performance, for the MasakhaPOS dataset. We find PLMs adapted for African languages to encode more linguistic information about target languages than massively multilingual PLMs. Our results reaffirm previous findings that token-level syntactic information concentrates in middle-to-last layers, while sentence-level semantic information is distributed across all layers. Through control tasks and probing baselines, we confirm that performance reflects the internal knowledge of PLMs rather than probe memorisation. Our study applies established interpretability techniques to African-language PLMs. In doing so, we highlight the internal mechanisms underlying the success of strategies like active learning and multilingual adaptation.

📄 PDF Abstract BibTeX arXiv:2505.10081

Code (0)

등록된 구현이 없습니다.

Tasks

Active LearningSentence

Similar Papers 제목 키워드 기반

LSR: Linguistic Safety Robustness Benchmark for Low-Resource West African Languages

2026-02-27 · Godwin Abuh Faruna arxiv

Safety alignment in large language models relies predominantly on English-language training data. When harmful intent is expressed in low-resource languages, refusal mechanisms that hold in English frequently fail to act…

Contextualising Levels of Language Resourcedness affecting Digital Processing of Text

2023-09-29 · C. Maria Keet, Langa Khumalo

Application domains such as digital humanities and tool like chatbots involve some form of processing natural language, from digitising hardcopies to speech generation. The language of the content is typically characteri…

AfroDigits: A Community-Driven Spoken Digit Dataset for African Languages

2023-03-22 · Chris Chinenye Emezue, Sanchit Gandhi, Lewis Tunstall, Abubakar Abid 외

The advancement of speech technologies has been remarkable, yet its integration with African languages remains limited due to the scarcity of African speech corpora. To address this issue, we present AfroDigits, a minima…

AfriMTE and AfriCOMET: Enhancing COMET to Embrace Under-resourced African Languages

2023-11-16 · Jiayi Wang, David Ifeoluwa Adelani, Sweta Agrawal, Marek Masiak 외

Despite the recent progress on scaling multilingual machine translation (MT) to several under-resourced African languages, accurately measuring this progress remains challenging, since evaluation is often performed on n-…

Machine Translation

Benchmarking Neural Machine Translation for Southern African Languages

2019-06-17 · WS 2019 8 · Laura Martinus, Jade Z. Abbott

Unlike major Western languages, most African languages are very low-resourced. Furthermore, the resources that do exist are often scattered and difficult to obtain and discover. As a result, the data and code for existin…

BenchmarkingMachine TranslationTranslation