paper-with-me

홈 › Papers

Enhancing expressivity transfer in textless speech-to-speech translation

2023-10-11 · Jarod Duret, Benjamin O'Brien, Yannick Estève, Titouan Parcollet

Textless speech-to-speech translation systems are rapidly advancing, thanks to the integration of self-supervised learning techniques. However, existing state-of-the-art systems fall short when it comes to capturing and transferring expressivity accurately across different languages. Expressivity plays a vital role in conveying emotions, nuances, and cultural subtleties, thereby enhancing communication across diverse languages. To address this issue this study presents a novel method that operates at the discrete speech unit level and leverages multilingual emotion embeddings to capture language-agnostic information. Specifically, we demonstrate how these embeddings can be used to effectively predict the pitch and duration of speech units in the target language. Through objective and subjective experiments conducted on a French-to-English translation task, our findings highlight the superior expressivity transfer achieved by our approach compared to current state-of-the-art systems.

📄 PDF Abstract BibTeX arXiv:2310.07279

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningSpeech-to-Speech TranslationTranslation

Similar Papers 제목 키워드 기반

Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation

2024-06-04 · Min-Jae Hwang, Ilia Kulikov, Benjamin Peloquin, Hongyu Gong 외

In this paper, we propose a textless acoustic model with a self-supervised distillation strategy for noise-robust expressive speech-to-speech translation (S2ST). Recently proposed expressive S2ST systems have achieved im…

Speech-to-Speech TranslationTranslation

Textless Unit-to-Unit training for Many-to-Many Multilingual Speech-to-Speech Translation

2023-08-03 · Minsu Kim, Jeongsoo Choi, Dahun Kim, Yong Man Ro

This paper proposes a textless training method for many-to-many multilingual speech-to-speech translation that can also benefit the transfer of pre-trained knowledge to text-based systems, text-to-speech synthesis and te…

DecoderQuantizationRepresentation LearningSpeech Synthesis+7

textless-lib: a Library for Textless Spoken Language Processing

2022-02-15 · NAACL (ACL) 2022 7 · Eugene Kharitonov, Jade Copet, Kushal Lakhotia, Tu Anh Nguyen 외

Textless spoken language processing research aims to extend the applicability of standard NLP toolset onto spoken language and languages with few or no textual resources. In this paper, we introduce textless-lib, a PyTor…

Resynthesis

Textless Speech-to-Speech Translation on Real Data

2021-12-15 · NAACL 2022 7 · Ann Lee, Hongyu Gong, Paul-Ambroise Duquenne, Holger Schwenk 외

We present a textless speech-to-speech translation (S2ST) system that can translate speech from one language into another language and can be built without the need of any text data. Different from existing work in the l…

Speech-to-Speech TranslationTranslation

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text

2025-09-09 · Peijin Xie, Shun Qian, Bingquan Liu, Dexin Wang 외 arxiv

Document images encapsulate a wealth of knowledge, while the portability of spoken queries enables broader and flexible application scenarios. Yet, no prior work has explored knowledge base question answering over visual…

Knowledge Base Question Answering