paper-with-me

홈 › Papers

Strategies for improving low resource speech to text translation relying on pre-trained ASR models

2023-05-31 · Santosh Kesiraju, Marek Sarvas, Tomas Pavlicek, Cecile Macaire, Alejandro Ciuba

This paper presents techniques and findings for improving the performance of low-resource speech to text translation (ST). We conducted experiments on both simulated and real-low resource setups, on language pairs English - Portuguese, and Tamasheq - French respectively. Using the encoder-decoder framework for ST, our results show that a multilingual automatic speech recognition system acts as a good initialization under low-resource scenarios. Furthermore, using the CTC as an additional objective for translation during training and decoding helps to reorder the internal representations and improves the final translation. Through our experiments, we try to identify various factors (initializations, objectives, and hyper-parameters) that contribute the most for improvements in low-resource setups. With only 300 hours of pre-training data, our model achieved 7.3 BLEU score on Tamasheq - French data, outperforming prior published works from IWSLT 2022 by 1.6 points.

📄 PDF Abstract BibTeX arXiv:2306.00208

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionDecoderspeech-recognitionSpeech RecognitionSpeech-to-TextSpeech-to-Text TranslationTranslation

Similar Papers 제목 키워드 기반

Textual Supervision for Visually Grounded Spoken Language Understanding

2020-10-06 · Findings of the Association for Computational Linguistics 2020 · Bertrand Higy, Desmond Elliott, Grzegorz Chrupała

Visually-grounded models of spoken language understanding extract semantic information directly from speech, without relying on transcriptions. This is useful for low-resource languages, where transcriptions can be expen…

Spoken Language Understanding

Leveraging unsupervised and weakly-supervised data to improve direct speech-to-speech translation

2022-03-24 · Ye Jia, Yifan Ding, Ankur Bapna, Colin Cherry 외

End-to-end speech-to-speech translation (S2ST) without relying on intermediate text representations is a rapidly emerging frontier of research. Recent works have demonstrated that the performance of such direct S2ST syst…

Representation LearningSpeech Representation LearningSpeech-to-Speech TranslationTranslation

AV-TranSpeech: Audio-Visual Robust Speech-to-Speech Translation

2023-05-24 · Rongjie Huang, Huadai Liu, Xize Cheng, Yi Ren 외

Direct speech-to-speech translation (S2ST) aims to convert speech from one language into another, and has demonstrated significant progress to date. Despite the recent success, current S2ST models still suffer from disti…

Speech-to-Speech TranslationTranslation

Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?

2025-10-03 · Oriol Pareras, Gerard I. Gállego, Federico Costa, Cristina España-Bonet 외 arxiv

Recent work on Speech-to-Text Translation (S2TT) has focused on LLM-based models, introducing the increasingly adopted Chain-of-Thought (CoT) prompting, where the model is guided to first transcribe the speech and then t…

Speech-to-Text TranslationSpeech Recognition

Gradient-Informed Training for Low-Resource Multilingual Speech Translation

2026-03-26 · Ruiyan Sun, Satoshi Nakamura arxiv

In low-resource multilingual speech-to-text translation, uniform architectural sharing across languages frequently introduces representation conflicts that impede convergence. This work proposes a principled methodology …

Speech-to-Text Translation