paper-with-me

홈 › Papers

WolBanking77: Wolof Banking Speech Intent Classification Dataset

2025-09-23 · Abdou Karim Kandji, Frédéric Precioso, Cheikh Ba, Samba Ndiaye, Augustin Ndione arxiv

Intent classification models have made a significant progress in recent years. However, previous studies primarily focus on high-resource language datasets, which results in a gap for low-resource languages and for regions with high rates of illiteracy, where languages are more spoken than read or written. This is the case in Senegal, for example, where Wolof is spoken by around 90\% of the population, while the national illiteracy rate remains at of 42\%. Wolof is actually spoken by more than 10 million people in West African region. To address these limitations, we introduce the Wolof Banking Speech Intent Classification Dataset (WolBanking77), for academic research in intent classification. WolBanking77 currently contains 9,791 text sentences in the banking domain and more than 4 hours of spoken sentences. Experiments on various baselines are conducted in this work, including text and voice state-of-the-art models. The results are very promising on this current dataset. In addition, this paper presents an in-depth examination of the dataset's contents. We report baseline F1-scores and word error rates metrics respectively on NLP and ASR models trained on WolBanking77 dataset and also comparisons between models. Dataset and code available at: https://github.com/abdoukarim/wolbanking77.

📄 PDF Abstract BibTeX arXiv:2509.19271

Code (0)

등록된 구현이 없습니다.

Tasks

Intent Classification

Similar Papers 제목 키워드 기반

Skit-S2I: An Indian Accented Speech to Intent dataset

2022-12-26 · Shangeth Rajaa, Swaraj Dalmia, Kumarmanas Nethil

Conventional conversation assistants extract text transcripts from the speech signal using automatic speech recognition (ASR) and then predict intent from the transcriptions. Using end-to-end spoken language understandin…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)intent-classificationIntent Classification+4

Cross-lingual Matryoshka Representation Learning across Speech and Text

2026-02-23 · Yaya Sy, Dioula Doucouré, Christophe Cerisara, Irina Illina arxiv

Speakers of under-represented languages face both a language barrier, as most online knowledge is in a few dominant languages, and a modality barrier, since information is largely text-based while many languages are prim…

Representation LearningIntent Detection

Implementation and Evaluation of an LFG-based Parser for Wolof

2020-05-01 · LREC 2020 5 · Cheikh M. Bamba Dione

This paper reports on a parsing system for Wolof based on the LFG formalism. The parser covers core constructions of Wolof, including noun classes, cleft, copula, causative and applicative sentences. It also deals with s…

Morphological Analysis

DarijaBanking: A New Resource for Overcoming Language Barriers in Banking Intent Detection for Moroccan Arabic Speakers

2024-05-26 · Abderrahman Skiredj, Ferdaous Azhari, Ismail Berrada, Saad Ezzini

Navigating the complexities of language diversity is a central challenge in developing robust natural language processing systems, especially in specialized domains like banking. The Moroccan Dialect (Darija) serves as t…

intent-classificationIntent ClassificationIntent DetectionLanguage Modeling+3

Speech Language Models for Under-Represented Languages: Insights from Wolof

2025-09-18 · Yaya Sy, Dioula Doucouré, Christophe Cerisara, Irina Illina arxiv

We present our journey in training a speech language model for Wolof, an underrepresented language spoken in West Africa, and share key insights. We first emphasize the importance of collecting large-scale, spontaneous, …

Speech Recognition