paper-with-me

Papers

Bemba Speech Translation: Exploring a Low-Resource African Language

2025-05-05 · Muhammad Hazim Al Farouq, Aman Kassahun Wassie, Yasmin Moslem

This paper describes our system submission to the International Conference on Spoken Language Translation (IWSLT 2025), low-resource languages track, namely for Bemba-to-English speech translation. We built cascaded speech translation systems based on Whisper and NLLB-200, and employed data augmentation techniques, such as back-translation. We investigate the effect of using synthetic data and discuss our experimental setup.

📄 PDF Abstract BibTeX arXiv:2505.02518

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationTranslation

Similar Papers 제목 키워드 기반

BIG-C: a Multimodal Multi-Purpose Dataset for Bemba

2023-05-26 · Claytone Sikasote, Eunice Mukonde, Md Mahfuz ibn Alam, Antonios Anastasopoulos

We present BIG-C (Bemba Image Grounded Conversations), a large multimodal dataset for Bemba. While Bemba is the most populous language of Zambia, it exhibits a dearth of resources which render the development of language…

Machine Translationspeech-recognitionSpeech RecognitionTranslation

BembaSpeech: A Speech Recognition Corpus for the Bemba Language

2021-02-09 · LREC 2022 6 · Claytone Sikasote, Antonios Anastasopoulos

We present a preprocessed, ready-to-use automatic speech recognition corpus, BembaSpeech, consisting over 24 hours of read speech in the Bemba language, a written but low-resourced language spoken by over 30% of the popu…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

KIT's Low-resource Speech Translation Systems for IWSLT2025: System Enhancement with Synthetic Data and Model Regularization

2025-05-26 · Zhaolin Li, Yining Liu, Danni Liu, Tuan Nam Nguyen 외

This paper presents KIT's submissions to the IWSLT 2025 low-resource track. We develop both cascaded systems, consisting of Automatic Speech Recognition (ASR) and Machine Translation (MT) models, and end-to-end (E2E) Spe…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+4

AfriVox-v2: A Domain-Verticalized Benchmark for In-the-Wild African Speech Recognition

2026-05-05 · Busayo Awobade, Gabrial Zencha Ashungafac, Tobi Olatunji arxiv

Recent large language models (LLMs) show strong speech recognition and translation capabilities for high-resource languages. However, African languages remain dramatically underrepresented in benchmarks, limiting their p…

Speech Recognition

OkwuGbé: End-to-End Speech Recognition for Fon and Igbo

2021-10-16 · ACL ARR October 2021 10 · Anonymous

Language is inherent and compulsory for human communication. Whether expressed in a written or spoken way, it ensures understanding between people of the same and different regions. With the growing awareness and effort …

Machine Translationspeech-recognitionSpeech Recognition