paper-with-me

홈 › Papers

Improving speech translation by fusing speech and text

2023-05-23 · Wenbiao Yin, Zhicheng Liu, Chengqi Zhao, Tao Wang, Jian Tong, Rong Ye

In speech translation, leveraging multimodal data to improve model performance and address limitations of individual modalities has shown significant effectiveness. In this paper, we harness the complementary strengths of speech and text, which are disparate modalities. We observe three levels of modality gap between them, denoted by Modal input representation, Modal semantic, and Modal hidden states. To tackle these gaps, we propose \textbf{F}use-\textbf{S}peech-\textbf{T}ext (\textbf{FST}), a cross-modal model which supports three distinct input modalities for translation: speech, text, and fused speech-text. We leverage multiple techniques for cross-modal alignment and conduct a comprehensive analysis to assess its impact on speech translation, machine translation, and fused speech-text translation. We evaluate FST on MuST-C, GigaST, and newstest benchmark. Experiments show that the proposed FST achieves an average 34.0 BLEU on MuST-C En$\rightarrow$De/Es/Fr (vs SOTA +1.1 BLEU). Further experiments demonstrate that FST does not degrade on MT task, as observed in prior works. Instead, it yields an average improvement of 3.2 BLEU over the pre-trained MT model.

📄 PDF Abstract BibTeX arXiv:2305.14042

Code (0)

등록된 구현이 없습니다.

Tasks

cross-modal alignmentMachine TranslationTranslation

Similar Papers 제목 키워드 기반

The Xiaomi Text-to-Text Simultaneous Speech Translation System for IWSLT 2022

2022-05-01 · IWSLT (ACL) 2022 5 · Bao Guo, Mengge Liu, Wen Zhang, Hexuan Chen 외

This system paper describes the Xiaomi Translation System for the IWSLT 2022 Simultaneous Speech Translation (noted as SST) shared task. We participate in the English-to-Mandarin Chinese Text-to-Text (noted as T2T) track…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationKnowledge Distillation+4

SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

2023-08-22 · Seamless Communication, Loïc Barrault, Yu-An Chung, Mariano Cora Meglioli 외

What does it take to create the Babel Fish, a tool that can help individuals translate speech between any two languages? While recent breakthroughs in text-based models have pushed machine translation coverage beyond 200…

Automatic Speech RecognitionMachine TranslationSpeech-to-Speech TranslationSpeech-to-Text+5

T-Modules: Translation Modules for Zero-Shot Cross-Modal Machine Translation

2022-05-24 · Paul-Ambroise Duquenne, Hongyu Gong, Benoît Sagot, Holger Schwenk

We present a new approach to perform zero-shot cross-modal transfer between speech and text for translation tasks. Multilingual speech and text are encoded in a joint fixed-size representation space. Then, we compare dif…

DecoderMachine Translationtext-to-speechText to Speech+2

Fused Acoustic and Text Encoding for Multimodal Bilingual Pretraining and Speech Translation

2021-02-10 · Renjie Zheng, Junkun Chen, Mingbo Ma, Liang Huang

Recently, representation learning for text and speech has successfully improved many language related tasks. However, all existing methods suffer from two limitations: (a) they only learn from one input modality, while a…

Language ModelingLanguage ModellingMachine TranslationRepresentation Learning+3

Textless Unit-to-Unit training for Many-to-Many Multilingual Speech-to-Speech Translation

2023-08-03 · Minsu Kim, Jeongsoo Choi, Dahun Kim, Yong Man Ro

This paper proposes a textless training method for many-to-many multilingual speech-to-speech translation that can also benefit the transfer of pre-trained knowledge to text-based systems, text-to-speech synthesis and te…

DecoderQuantizationRepresentation LearningSpeech Synthesis+7