paper-with-me

홈 › Papers

Yankari: A Monolingual Yoruba Dataset

2024-12-04 · Maro Akpobi

This paper presents Yankari, a large-scale monolingual dataset for the Yoruba language, aimed at addressing the critical gap in Natural Language Processing (NLP) resources for this important West African language. Despite being spoken by over 30 million people, Yoruba has been severely underrepresented in NLP research and applications. We detail our methodology for creating this dataset, which includes careful source selection, automated quality control, and rigorous data cleaning processes. The Yankari dataset comprises 51,407 documents from 13 diverse sources, totaling over 30 million tokens. Our approach focuses on ethical data collection practices, avoiding problematic sources and addressing issues prevalent in existing datasets. We provide thorough automated evaluations of the dataset, demonstrating its quality compared to existing resources. The Yankari dataset represents a significant advancement in Yoruba language resources, providing a foundation for developing more accurate NLP models, supporting comparative linguistic studies, and contributing to the digital accessibility of the Yoruba language.

📄 PDF Abstract BibTeX arXiv:2412.03334

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

When Does Language Transfer Help? Sequential Fine-Tuning for Cross-Lingual Euphemism Detection

2025-08-15 · Julia Sammartino, Libby Barak, Jing Peng, Anna Feldman arxiv

Euphemisms are culturally variable and often ambiguous, posing challenges for language models, especially in low-resource settings. This paper investigates how cross-lingual transfer via sequential fine-tuning affects eu…

Cross-Lingual Transfer

HausaNLP at SemEval-2023 Task 12: Leveraging African Low Resource TweetData for Sentiment Analysis

2023-04-26 · Saheed Abdullahi Salahudeen, Falalu Ibrahim Lawan, Ahmad Mustapha Wali, Amina Abubakar Imam 외

We present the findings of SemEval-2023 Task 12, a shared task on sentiment analysis for low-resource African languages using Twitter dataset. The task featured three subtasks; subtask A is monolingual sentiment classifi…

Sentiment AnalysisSentiment ClassificationTwitter Sentiment AnalysisZero-shot Sentiment Classification

A Situational Speech Synthesizer for Yoruba: System Design, Phonological Rule Architecture, and Orthographic Extensions for Contour

2026-07-17 · Kola Tubosun, Adedayo Oluokun, Hafiz Adewuyi, Dadepo Aderemi arxiv

We present TTSYoruba, a rule-based concatenative diphone speech synthesizer for Yoruba, deployed at online as part of the YorubaName.com open dictionary of Yoruba personal names. The system takes tone-marked Yoruba text …

Visually Grounded Speech Models for Low-resource Languages and Cognitive Modelling

2024-09-03 · Leanne Nortje

This dissertation examines visually grounded speech (VGS) models that learn from unlabelled speech paired with images. It focuses on applications for low-resource languages and understanding human language acquisition. W…

Few-Shot LearningLanguage Acquisition

The Challenge of Diacritics in Yoruba Embeddings

2020-11-15 · Tosin P. Adewumi, Foteini Liwicki, Marcus Liwicki

The major contributions of this work include the empirical establishment of a better performance for Yoruba embeddings from undiacritized (normalized) dataset and provision of new analogy sets for evaluation. The Yoruba …