paper-with-me

홈 › Papers

Self-Supervised Representations Improve End-to-End Speech Translation

2020-06-22 · Anne Wu, Changhan Wang, Juan Pino, Jiatao Gu

End-to-end speech-to-text translation can provide a simpler and smaller system but is facing the challenge of data scarcity. Pre-training methods can leverage unlabeled data and have been shown to be effective on data-scarce settings. In this work, we explore whether self-supervised pre-trained speech representations can benefit the speech translation task in both high- and low-resource settings, whether they can transfer well to other languages, and whether they can be effectively combined with other common methods that help improve low-resource end-to-end speech translation such as using a pre-trained high-resource speech recognition system. We demonstrate that self-supervised pre-trained features can consistently improve the translation performance, and cross-lingual transfer allows to extend to a variety of languages without or with little tuning.

📄 PDF Abstract BibTeX arXiv:2006.12124

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Lingual Transferspeech-recognitionSpeech RecognitionSpeech-to-TextSpeech-to-Text TranslationTranslation

Similar Papers 제목 키워드 기반

Enhanced Direct Speech-to-Speech Translation Using Self-supervised Pre-training and Data Augmentation

2022-04-06 · Sravya Popuri, Peng-Jen Chen, Changhan Wang, Juan Pino 외

Direct speech-to-speech translation (S2ST) models suffer from data scarcity issues as there exists little parallel S2ST data, compared to the amount of data available for conventional cascaded systems that consist of aut…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationDecoder+9

Speech-to-Speech Translation with Discrete-Unit-Based Style Transfer

2023-09-14 · Yongqi Wang, Jionghao Bai, Rongjie Huang, RuiQi Li 외

Direct speech-to-speech translation (S2ST) with discrete self-supervised representations has achieved remarkable accuracy, but is unable to preserve the speaker timbre of the source speech. Meanwhile, the scarcity of hig…

In-Context LearningLanguage ModelingLanguage ModellingSpeech-to-Speech Translation+2

SparQLe: Speech Queries to Text Translation Through LLMs

2025-02-13 · Amirbek Djanibekov, Hanan Aldarmaki

With the growing influence of Large Language Models (LLMs), there is increasing interest in integrating speech representations with them to enable more seamless multi-modal processing and speech understanding. This study…

Speech-to-TextSpeech-to-Text TranslationTranslation

Unified Speech-Text Pre-training for Speech Translation and Recognition

2022-04-11 · ACL 2022 5 · Yun Tang, Hongyu Gong, Ning Dong, Changhan Wang 외

We describe a method to jointly pre-train speech and text in an encoder-decoder modeling framework for speech translation and recognition. The proposed method incorporates four self-supervised and supervised subtasks for…

Decoderspeech-recognitionSpeech RecognitionTranslation

MAESTRO: Matched Speech Text Representations through Modality Matching

2022-04-07 · Zhehuai Chen, Yu Zhang, Andrew Rosenberg, Bhuvana Ramabhadran 외

We present Maestro, a self-supervised training method to unify representations learnt from speech and text modalities. Self-supervised learning from speech signals aims to learn the latent structure inherent in the signa…

Language ModellingSelf-Supervised Learningspeech-recognitionSpeech Recognition+1