Speech Translation Refinement using Large Language Models
Recent advancements in large language models (LLMs) have demonstrated their remarkable capabilities across various language tasks. Inspired by the success of text-to-text translation refinement, this paper investigates how LLMs can improve the performance of speech translation by introducing a joint refinement process. Through the joint refinement of speech translation (ST) and automatic speech recognition (ASR) transcription via LLMs, the performance of the ST model is significantly improved in both training-free in-context learning and parameter-efficient fine-tuning scenarios. Additionally, we explore the effect of document-level context on refinement under the context-aware fine-tuning scenario. Experimental results on the MuST-C and CoVoST 2 datasets, which include seven translation tasks, demonstrate the effectiveness of the proposed approach using several popular LLMs including GPT-3.5-turbo, LLaMA3-8B, and Mistral-12B. Further analysis further suggests that jointly refining both transcription and translation yields better performance compared to refining translation alone. Meanwhile, incorporating document-level context significantly enhances refinement performance. We release our code and datasets on GitHub.
Code (1)
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)In-Context Learningparameter-efficient fine-tuningspeech-recognitionSpeech RecognitionTranslationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
KIT's Offline Speech Translation and Instruction Following Submission for IWSLT 2025
The scope of the International Workshop on Spoken Language Translation (IWSLT) has recently broadened beyond traditional Speech Translation (ST) to encompass a wider array of tasks, including Speech Question Answering an…
Automatic Speech RecognitionInstruction FollowingQuestion Answeringspeech-recognition+2DoCIA: An Online Document-Level Context Incorporation Agent for Speech Translation
Document-level context is crucial for handling discourse challenges in text-to-text document-level machine translation (MT). Despite the increased discourse challenges introduced by noise from automatic speech recognitio…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Document Level Machine TranslationLanguage Modeling+7Unsupervised Cross-Modal Alignment of Speech and Text Embedding Spaces
Recent research has shown that word embedding spaces learned from text corpora of different languages can be aligned without any parallel data supervision. Inspired by the success in unsupervised cross-lingual word embed…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Cross-Lingual Word Embeddingscross-modal alignment+6CoVoST 2 and Massively Multilingual Speech-to-Text Translation
Speech translation has recently become an increasingly popular topic of research, partly due to the development of benchmark datasets. Nevertheless, current datasets cover a limited number of languages. With the aim to f…
Machine Translationspeech-recognitionSpeech RecognitionSpeech-to-Text+2Steering LLMs toward Korean Local Speech: Iterative Refinement Framework for Faithful Dialect Translation
Standard-to-dialect machine translation remains challenging due to a persistent dialect gap in large language models and evaluation distortions inherent in n-gram metrics, which favor source copying over authentic dialec…
Machine Translation