paper-with-me

홈 › Papers

FFSTC: Fongbe to French Speech Translation Corpus

2024-03-08 · D. Fortune Kponou, Frejus A. A. Laleye, Eugene C. Ezin

In this paper, we introduce the Fongbe to French Speech Translation Corpus (FFSTC) for the first time. This corpus encompasses approximately 31 hours of collected Fongbe language content, featuring both French transcriptions and corresponding Fongbe voice recordings. FFSTC represents a comprehensive dataset compiled through various collection methods and the efforts of dedicated individuals. Furthermore, we conduct baseline experiments using Fairseq's transformer_s and conformer models to evaluate data quality and validity. Our results indicate a score of 8.96 for the transformer_s model and 8.14 for the conformer model, establishing a baseline for the FFSTC corpus.

📄 PDF Abstract BibTeX arXiv:2403.05488

Code (0)

등록된 구현이 없습니다.

Tasks

Translation

Similar Papers 제목 키워드 기반

When Does Data Augmentation Help? Evaluating LLM and Back-Translation Methods for Hausa and Fongbe NLP

2026-04-14 · Mahounan Pericles Adjovi, Roald Eiselen, Prasenjit Mitra arxiv

Data scarcity limits NLP development for low-resource African languages. We evaluate two data augmentation methods -- LLM-based generation (Gemini 2.5 Flash) and back-translation (NLLB-200) -- for Hausa and Fongbe, two W…

Data AugmentationPOS Tagging

From Speech to Text Corpora: Evaluating ASR-Based Data Acquisition for Low-Resource Fongbe and Hausa

2026-06-20 · Mahounan Pericles Adjovi, Victor Olufemi, Roald Eiselen, Prasenjit Mitra arxiv

Low-resource African languages lack text corpora needed for language model training. We investigate whether ASR pipelines can extend text resources for two typologically distinct West African languages: Fongbe (tonal, di…

Microsoft Speech Language Translation (MSLT) Corpus: The IWSLT 2016 release for English, French and German

2016-12-01 · IWSLT 2016 12 · Christian Federmann, William D. Lewis

We describe the Microsoft Speech Language Translation (MSLT) corpus, which was created in order to evaluate end-to-end conversational speech translation quality. The corpus was created from actual conversations over Skyp…

Machine Translationspeech-recognitionSpeech RecognitionTranslation

First Automatic Fongbe Continuous Speech Recognition System: Development of Acoustic Models and Language Models

2017-01-21 · Fréjus Laleye, Laurent Besacier, Eugène Ezin, Cina Motamed.

This paper reports our efforts toward an ASR system for a new under-resourced language (Fongbe). The aim of this work is to build acoustic models and language models for continuous speech decoding in Fongbe. The problem …

Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Augmenting Librispeech with French Translations: A Multimodal Corpus for Direct Speech Translation Evaluation

2018-02-09 · LREC 2018 5 · Ali Can Kocabiyikoglu, Laurent Besacier, Olivier Kraif

Recent works in spoken language translation (SLT) have attempted to build end-to-end speech-to-text translation without using source language transcription during learning or decoding. However, while large quantities of …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine TranslationSentence+5