paper-with-me

Papers

Simple and Effective Unsupervised Speech Translation

2022-10-18 · Changhan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov, Yun Tang, Wei-Ning Hsu, Michael Auli, Juan Pino

The amount of labeled data to train models for speech tasks is limited for most languages, however, the data scarcity is exacerbated for speech translation which requires labeled data covering two different languages. To address this issue, we study a simple and effective approach to build speech translation systems without labeled data by leveraging recent advances in unsupervised speech recognition, machine translation and speech synthesis, either in a pipeline approach, or to generate pseudo-labels for training end-to-end speech translation models. Furthermore, we present an unsupervised domain adaptation technique for pre-trained speech models which improves the performance of downstream unsupervised speech recognition, especially for low-resource settings. Experiments show that unsupervised speech-to-text translation outperforms the previous unsupervised state of the art by 3.2 BLEU on the Libri-Trans benchmark, on CoVoST 2, our best systems outperform the best supervised end-to-end models (without pre-training) from only two years ago by an average of 5.0 BLEU over five X-En directions. We also report competitive results on MuST-C and CVSS benchmarks.

📄 PDF Abstract BibTeX arXiv:2210.10191

Code (0)

등록된 구현이 없습니다.

Tasks

Domain AdaptationMachine Translationspeech-recognitionSpeech RecognitionSpeech SynthesisSpeech-to-TextSpeech-to-Text TranslationTranslationUnsupervised Domain AdaptationUnsupervised Speech Recognition

Similar Papers 제목 키워드 기반

Towards speech-to-text translation without speech recognition

2017-02-13 · EACL 2017 4 · Sameer Bansal, Herman Kamper, Adam Lopez, Sharon Goldwater

We explore the problem of translating speech to text in low-resource scenarios where neither automatic speech recognition (ASR) nor machine translation (MT) are available, but we have training data in the form of audio p…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+4

Simple and Effective Unsupervised Speech Synthesis

2022-04-06 · Alexander H. Liu, Cheng-I Jeff Lai, Wei-Ning Hsu, Michael Auli 외

We introduce the first unsupervised speech synthesis system based on a simple, yet effective recipe. The framework leverages recent work in unsupervised speech recognition as well as existing neural-based speech synthesi…

speech-recognitionSpeech RecognitionSpeech SynthesisUnsupervised Speech Recognition

Improving Cascaded Unsupervised Speech Translation with Denoising Back-translation

2023-05-12 · Yu-Kuan Fu, Liang-Hsuan Tseng, Jiatong Shi, Chen-An Li 외

Most of the speech translation models heavily rely on parallel data, which is hard to collect especially for low-resource languages. To tackle this issue, we propose to build a cascaded speech translation system without …

DenoisingMachine TranslationTranslation

Translatotron 3: Speech to Speech Translation with Monolingual Data

2023-05-27 · Eliya Nachmani, Alon Levkovitch, Yifan Ding, Chulayuth Asawaroengchai 외

This paper presents Translatotron 3, a novel approach to unsupervised direct speech-to-speech translation from monolingual speech-text datasets by combining masked autoencoder, unsupervised embedding mapping, and back-tr…

Speech-to-Speech TranslationTranslation

Leveraging unsupervised and weakly-supervised data to improve direct speech-to-speech translation

2022-03-24 · Ye Jia, Yifan Ding, Ankur Bapna, Colin Cherry 외

End-to-end speech-to-speech translation (S2ST) without relying on intermediate text representations is a rapidly emerging frontier of research. Recent works have demonstrated that the performance of such direct S2ST syst…

Representation LearningSpeech Representation LearningSpeech-to-Speech TranslationTranslation