paper-with-me

홈 › Papers

UWSpeech: Speech to Speech Translation for Unwritten Languages

2020-06-14 · Chen Zhang, Xu Tan, Yi Ren, Tao Qin, Ke-jun Zhang, Tie-Yan Liu

Existing speech to speech translation systems heavily rely on the text of target language: they usually translate source language either to target text and then synthesize target speech from text, or directly to target speech with target text for auxiliary training. However, those methods cannot be applied to unwritten target languages, which have no written text or phoneme available. In this paper, we develop a translation system for unwritten languages, named as UWSpeech, which converts target unwritten speech into discrete tokens with a converter, and then translates source-language speech into target discrete tokens with a translator, and finally synthesizes target speech from target discrete tokens with an inverter. We propose a method called XL-VAE, which enhances vector quantized variational autoencoder (VQ-VAE) with cross-lingual (XL) speech recognition, to train the converter and inverter of UWSpeech jointly. Experiments on Fisher Spanish-English conversation translation dataset show that UWSpeech outperforms direct translation and VQ-VAE baseline by about 16 and 10 BLEU points respectively, which demonstrate the advantages and potentials of UWSpeech.

📄 PDF Abstract BibTeX arXiv:2006.07926

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech RecognitionSpeech-to-Speech TranslationTranslation

Methods 이 논문이 사용한 방법론

VQ-VAE VQ-VAE is a type of variational autoencoder that uses vector quantisation to obtain a discrete latent representation. It differs from…
Solana Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

Assessing Evaluation Metrics for Speech-to-Speech Translation

2021-10-26 · Elizabeth Salesky, Julian Mäder, Severin Klinger

Speech-to-speech translation combines machine translation with speech synthesis, introducing evaluation challenges not present in either task alone. How to automatically evaluate speech-to-speech translation is an open q…

Machine TranslationOpen-Ended Question AnsweringSpeech SynthesisSpeech-to-Speech Translation+1

PolyVoice: Language Models for Speech to Speech Translation

2023-06-05 · Qianqian Dong, Zhiying Huang, Qiao Tian, Chen Xu 외

We propose PolyVoice, a language model-based framework for speech-to-speech translation (S2ST) system. Our framework consists of two language models: a translation language model and a speech synthesis language model. We…

Language ModelingLanguage ModellingSpeech SynthesisSpeech-to-Speech Translation+1

Direct speech-to-speech translation with discrete units

2021-07-12 · ACL 2022 5 · Ann Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu 외

We present a direct speech-to-speech translation (S2ST) model that translates speech from one language to speech in another language without relying on intermediate text generation. We tackle the problem by first applyin…

Speech-to-Speech TranslationText GenerationTranslation

Listen and Translate: A Proof of Concept for End-to-End Speech-to-Text Translation

2016-12-06 · Alexandre Berard, Olivier Pietquin, Christophe Servan, Laurent Besacier

This paper proposes a first attempt to build an end-to-end speech-to-text translation system, which does not use source language transcription during learning or decoding. We propose a model for direct speech-to-text tra…

Speech-to-TextSpeech-to-Text TranslationTranslation

Speech-to-Speech Translation For A Real-world Unwritten Language

2022-11-11 · arXiv 2022 10 · Peng-Jen Chen, Kevin Tran, Yilin Yang, Jingfei Du 외

We study speech-to-speech translation (S2ST) that translates speech from one language into another language and focuses on building systems to support languages without standard text writing systems. We use English-Taiwa…

Speech-to-Speech TranslationTranslation