paper-with-me

홈 › Papers

Speech Resources in the Tamasheq Language

2022-01-13 · LREC 2022 6 · Marcely Zanon Boito, Fethi Bougares, Florentin Barbier, Souhir Gahbiche, Loïc Barrault, Mickael Rouvier, Yannick Estève

In this paper we present two datasets for Tamasheq, a developing language mainly spoken in Mali and Niger. These two datasets were made available for the IWSLT 2022 low-resource speech translation track, and they consist of collections of radio recordings from daily broadcast news in Niger (Studio Kalangou) and Mali (Studio Tamani). We share (i) a massive amount of unlabeled audio data (671 hours) in five languages: French from Niger, Fulfulde, Hausa, Tamasheq and Zarma, and (ii) a smaller 17 hours parallel corpus of audio recordings in Tamasheq, with utterance-level translations in the French language. All this data is shared under the Creative Commons BY-NC-ND 3.0 license. We hope these resources will inspire the speech community to develop and benchmark models using the Tamasheq language.

📄 PDF Abstract BibTeX arXiv:2201.05051

Code (1)

mzboito/iwslt2022_tamasheq_data 공식 구현

Tasks

Translation

Similar Papers 제목 키워드 기반

ON-TRAC Consortium Systems for the IWSLT 2022 Dialect and Low-resource Speech Translation Tasks

2022-05-04 · IWSLT (ACL) 2022 5 · Marcely Zanon Boito, John Ortega, Hugo Riguidel, Antoine Laurent 외

This paper describes the ON-TRAC Consortium translation systems developed for two challenge tracks featured in the Evaluation Campaign of IWSLT 2022: low-resource and dialect speech translation. For the Tunisian Arabic-E…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+3

Strategies for improving low resource speech to text translation relying on pre-trained ASR models

2023-05-31 · Santosh Kesiraju, Marek Sarvas, Tomas Pavlicek, Cecile Macaire 외

This paper presents techniques and findings for improving the performance of low-resource speech to text translation (ST). We conducted experiments on both simulated and real-low resource setups, on language pairs Englis…

Automatic Speech RecognitionDecoderspeech-recognitionSpeech Recognition+3

NAVER LABS Europe's Multilingual Speech Translation Systems for the IWSLT 2023 Low-Resource Track

2023-06-13 · Edward Gow-Smith, Alexandre Berard, Marcely Zanon Boito, Ioan Calapodescu

This paper presents NAVER LABS Europe's systems for Tamasheq-French and Quechua-Spanish speech translation in the IWSLT 2023 Low-Resource track. Our work attempts to maximize translation quality in low-resource settings …

Translation

Tools and resources for Romanian text-to-speech and speech-to-text applications

2018-02-15 · Tiberiu Boros, Stefan Daniel Dumitrescu, Vasile Pais

In this paper we introduce a set of resources and tools aimed at providing support for natural language processing, text-to-speech synthesis and speech recognition for Romanian. While the tools are general purpose and ca…

speech-recognitionSpeech RecognitionSpeech SynthesisSpeech-to-Text+3

Hidden Resources ― Strategies to Acquire and Exploit Potential Spoken Language Resources in National Archives

2016-05-01 · LREC 2016 5 · Jens Edlund, Joakim Gustafson

In 2014, the Swedish government tasked a Swedish agency, The Swedish Post and Telecom Authority (PTS), with investigating how to best create and populate an infrastructure for spoken language resources (Ref N2014/2840/IT…

Position