Enhancing End-to-End Conversational Speech Translation Through Target Language Context Utilization
Incorporating longer context has been shown to benefit machine translation, but the inclusion of context in end-to-end speech translation (E2E-ST) remains under-studied. To bridge this gap, we introduce target language context in E2E-ST, enhancing coherence and overcoming memory constraints of extended audio segments. Additionally, we propose context dropout to ensure robustness to the absence of context, and further improve performance by adding speaker information. Our proposed contextual E2E-ST outperforms the isolated utterance-based E2E-ST approach. Lastly, we demonstrate that in conversational speech, contextual information primarily contributes to capturing context style, as well as resolving anaphora and named entities.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationTranslationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Advancing Speech Translation: A Corpus of Mandarin-English Conversational Telephone Speech
This paper introduces a set of English translations for a 123-hour subset of the CallHome Mandarin Chinese data and the HKUST Mandarin Telephone Speech data for the task of speech translation. Paired source-language spee…
TranslationEnhancing expressivity transfer in textless speech-to-speech translation
Textless speech-to-speech translation systems are rapidly advancing, thanks to the integration of self-supervised learning techniques. However, existing state-of-the-art systems fall short when it comes to capturing and …
Self-Supervised LearningSpeech-to-Speech TranslationTranslationFluent Translations from Disfluent Speech in End-to-End Speech Translation
Spoken language translation applications for speech suffer due to conversational speech phenomena, particularly the presence of disfluencies. With the rise of end-to-end speech translation models, processing steps such a…
Machine Translationspeech-recognitionSpeech RecognitionTranslationCross-Lingual Conversational Speech Summarization with Large Language Models
Cross-lingual conversational speech summarization is an important problem, but suffers from a dearth of resources. While transcriptions exist for a number of languages, translated conversational speech is rare and datase…
Machine Translationspeech-recognitionSpeech RecognitionTranslationGenerating Fluent Translations from Disfluent Text Without Access to Fluent References: IIT Bombay@IWSLT2020
Machine translation systems perform reasonably well when the input is well-formed speech or text. Conversational speech is spontaneous and inherently consists of many disfluencies. Producing fluent translations of disflu…
DenoisingMachine TranslationTranslation