paper-with-me

홈 › Papers

TranSentence: Speech-to-speech Translation via Language-agnostic Sentence-level Speech Encoding without Language-parallel Data

2024-01-17 · Seung-bin Kim, Sang-Hoon Lee, Seong-Whan Lee

Although there has been significant advancement in the field of speech-to-speech translation, conventional models still require language-parallel speech data between the source and target languages for training. In this paper, we introduce TranSentence, a novel speech-to-speech translation without language-parallel speech data. To achieve this, we first adopt a language-agnostic sentence-level speech encoding that captures the semantic information of speech, irrespective of language. We then train our model to generate speech based on the encoded embedding obtained from a language-agnostic sentence-level speech encoder that is pre-trained with various languages. With this method, despite training exclusively on the target language's monolingual data, we can generate target language speech in the inference stage using language-agnostic speech embedding from the source language speech. Furthermore, we extend TranSentence to multilingual speech-to-speech translation. The experimental results demonstrate that TranSentence is superior to other models.

📄 PDF Abstract BibTeX arXiv:2401.12992

Code (0)

등록된 구현이 없습니다.

Tasks

SentenceSpeech-to-Speech TranslationTranslation

Similar Papers 제목 키워드 기반

Code-Switching without Switching: Language Agnostic End-to-End Speech Translation

2022-10-04 · Christian Huber, Enes Yavuz Ugan, Alexander Waibel

We propose a) a Language Agnostic end-to-end Speech Translation model (LAST), and b) a data augmentation strategy to increase code-switching (CS) performance. With increasing globalization, multiple languages are increas…

Data Augmentationspeech-recognitionSpeech RecognitionTranslation

Does Translation-Enhanced Speech Encoder Pre-training Affect Speech LLMs?

2026-06-24 · Tomoya Mizumoto, Yusuke Fujita arxiv

Connecting a pre-trained speech encoder to a Large Language Model (LLM) is the standard architecture for building Speech LLMs. However, a structural misalignment exists between the encoder and the LLM. Unlike encoders ba…

Speech Recognition

Representation Purification for End-to-End Speech Translation

2024-12-05 · Chengwei Zhang, Yue Zhou, Rui Zhao, Yidong Chen 외

Speech-to-text translation (ST) is a cross-modal task that involves converting spoken language into text in a different language. Previous research primarily focused on enhancing speech translation by facilitating knowle…

Machine TranslationRhythmSpeech-to-TextSpeech-to-Text Translation+2

LAMASSU: Streaming Language-Agnostic Multilingual Speech Recognition and Translation Using Neural Transducers

2022-11-05 · Peidong Wang, Eric Sun, Jian Xue, Yu Wu 외

Automatic speech recognition (ASR) and speech translation (ST) can both use neural transducers as the model structure. It is thus possible to use a single transducer model to perform both tasks. In real-world application…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Identificationspeech-recognition+3

Enhancing expressivity transfer in textless speech-to-speech translation

2023-10-11 · Jarod Duret, Benjamin O'Brien, Yannick Estève, Titouan Parcollet

Textless speech-to-speech translation systems are rapidly advancing, thanks to the integration of self-supervised learning techniques. However, existing state-of-the-art systems fall short when it comes to capturing and …

Self-Supervised LearningSpeech-to-Speech TranslationTranslation