paper-with-me

홈 › Papers

First Automatic Fongbe Continuous Speech Recognition System: Development of Acoustic Models and Language Models

2017-01-21 · Fréjus Laleye, Laurent Besacier, Eugène Ezin, Cina Motamed.

This paper reports our efforts toward an ASR system for a new under-resourced language (Fongbe). The aim of this work is to build acoustic models and language models for continuous speech decoding in Fongbe. The problem encountered with Fongbe (an African language spoken especially in Benin, Togo, and Nigeria) is that it does not have any language resources for an ASR system. As part of this work, we have first collected Fongbe text and speech corpora that are described in the following sections. Acoustic modeling has been worked out at a graphemic level and language modeling has provided two language models for performance comparison purposes. We also performed a vowel simplification by removing tones diacritics in order to investigate their impact on the language models.

📄 PDF Abstract BibTeX

Code (1)

laleye/ALFFA_PUBLIC

Tasks

Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

FFSTC: Fongbe to French Speech Translation Corpus

2024-03-08 · D. Fortune Kponou, Frejus A. A. Laleye, Eugene C. Ezin

In this paper, we introduce the Fongbe to French Speech Translation Corpus (FFSTC) for the first time. This corpus encompasses approximately 31 hours of collected Fongbe language content, featuring both French transcript…

Translation

When Does Data Augmentation Help? Evaluating LLM and Back-Translation Methods for Hausa and Fongbe NLP

2026-04-14 · Mahounan Pericles Adjovi, Roald Eiselen, Prasenjit Mitra arxiv

Data scarcity limits NLP development for low-resource African languages. We evaluate two data augmentation methods -- LLM-based generation (Gemini 2.5 Flash) and back-translation (NLLB-200) -- for Hausa and Fongbe, two W…

Data AugmentationPOS Tagging

A Survey of Text and Speech Resources for Hausa and Fongbe: Availability, Quality, and Gaps for NLP Development

2026-04-13 · Mahounan Pericles Adjovi, Victor Olufemi, Roald Eiselen, Prasenjit Mitra arxiv

This survey provides a comprehensive catalog of publicly available text and speech resources for two West African languages: Hausa, an Afroasiatic language with approximately 80-100 million speakers, and Fongbe, a Niger-…

POS Tagging

Evaluating Large Language Models for Hausa and Fongbe Machine Translation: Benchmarks, Failures, and Metric Reliability

2026-06-20 · Mahounan Pericles Adjovi, Roald Eiselen, Prasenjit Mitra arxiv

We investigate the translation quality of current large language models (LLMs) for English-to-Hausa and English-to-Fongbe - two typologically distinct West African languages from the Afroasiatic and Niger-Congo families …

Machine Translation

From Speech to Text Corpora: Evaluating ASR-Based Data Acquisition for Low-Resource Fongbe and Hausa

2026-06-20 · Mahounan Pericles Adjovi, Victor Olufemi, Roald Eiselen, Prasenjit Mitra arxiv

Low-resource African languages lack text corpora needed for language model training. We investigate whether ASR pipelines can extend text resources for two typologically distinct West African languages: Fongbe (tonal, di…