Strategies for improving low resource speech to text translation relying on pre-trained ASR models
This paper presents techniques and findings for improving the performance of low-resource speech to text translation (ST). We conducted experiments on both simulated and real-low resource setups, on language pairs English - Portuguese, and Tamasheq - French respectively. Using the encoder-decoder framework for ST, our results show that a multilingual automatic speech recognition system acts as a good initialization under low-resource scenarios. Furthermore, using the CTC as an additional objective for translation during training and decoding helps to reorder the internal representations and improves the final translation. Through our experiments, we try to identify various factors (initializations, objectives, and hyper-parameters) that contribute the most for improvements in low-resource setups. With only 300 hours of pre-training data, our model achieved 7.3 BLEU score on Tamasheq - French data, outperforming prior published works from IWSLT 2022 by 1.6 points.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionDecoderspeech-recognitionSpeech RecognitionSpeech-to-TextSpeech-to-Text TranslationTranslationSimilar Papers 제목 키워드 기반
Textual Supervision for Visually Grounded Spoken Language Understanding
Visually-grounded models of spoken language understanding extract semantic information directly from speech, without relying on transcriptions. This is useful for low-resource languages, where transcriptions can be expen…
Spoken Language UnderstandingLeveraging unsupervised and weakly-supervised data to improve direct speech-to-speech translation
End-to-end speech-to-speech translation (S2ST) without relying on intermediate text representations is a rapidly emerging frontier of research. Recent works have demonstrated that the performance of such direct S2ST syst…
Representation LearningSpeech Representation LearningSpeech-to-Speech TranslationTranslationAV-TranSpeech: Audio-Visual Robust Speech-to-Speech Translation
Direct speech-to-speech translation (S2ST) aims to convert speech from one language into another, and has demonstrated significant progress to date. Despite the recent success, current S2ST models still suffer from disti…
Speech-to-Speech TranslationTranslationRevisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?
Recent work on Speech-to-Text Translation (S2TT) has focused on LLM-based models, introducing the increasingly adopted Chain-of-Thought (CoT) prompting, where the model is guided to first transcribe the speech and then t…
Speech-to-Text TranslationSpeech RecognitionGradient-Informed Training for Low-Resource Multilingual Speech Translation
In low-resource multilingual speech-to-text translation, uniform architectural sharing across languages frequently introduces representation conflicts that impede convergence. This work proposes a principled methodology …
Speech-to-Text Translation