paper-with-me

홈 › Papers

SeamlessExpressiveLM: Speech Language Model for Expressive Speech-to-Speech Translation with Chain-of-Thought

2024-05-30 · Hongyu Gong, Bandhav Veluri

Expressive speech-to-speech translation (S2ST) is a key research topic in seamless communication, which focuses on the preservation of semantics and speaker vocal style in translated speech. Early works synthesized speaker style aligned speech in order to directly learn the mapping from speech to target speech spectrogram. Without reliance on style aligned data, recent studies leverage the advances of language modeling (LM) and build cascaded LMs on semantic and acoustic tokens. This work proposes SeamlessExpressiveLM, a single speech language model for expressive S2ST. We decompose the complex source-to-target speech mapping into intermediate generation steps with chain-of-thought prompting. The model is first guided to translate target semantic content and then transfer the speaker style to multi-stream acoustic units. Evaluated on Spanish-to-English and Hungarian-to-English translations, SeamlessExpressiveLM outperforms cascaded LMs in both semantic quality and style transfer, meanwhile achieving better parameter efficiency.

📄 PDF Abstract BibTeX arXiv:2405.20410

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingSpeech-to-Speech TranslationStyle Transfer

Similar Papers 제목 키워드 기반

SpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language Description

2024-08-24 · Zeyu Jin, Jia Jia, Qixin Wang, Kehan Li 외

Speech-language multi-modal learning presents a significant challenge due to the fine nuanced information inherent in speech styles. Therefore, a large-scale dataset providing elaborate comprehension of speech style is u…

DescriptiveSpeech SynthesisTAG

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation

2025-08-22 · Weiting Tan, Jiachen Lian, Hirofumi Inaguma, Paden Tomasello 외 arxiv

We present an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model. We explore multiple visual encoders and multimodal fusion…

Emotion Recognition

Expressive Text-to-Speech using Style Tag

2021-04-01 · Minchan Kim, Sung Jun Cheon, Byoung Jin Choi, Jong Jin Kim 외

As recent text-to-speech (TTS) systems have been rapidly improved in speech quality and generation speed, many researchers now focus on a more challenging issue: expressive TTS. To control speaking styles, existing expre…

Language ModelingLanguage ModellingTAGtext-to-speech+1

UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice

2025-09-25 · Sitong Cheng, Weizhen Bian, Xinsheng Wang, Ruibin Yuan 외 arxiv

The ultimate goal of expressive speech-to-speech translation (S2ST) is to accurately translate spoken content while preserving the speaker identity and emotional style. However, progress in this field is largely hindered…

Speech-to-Speech TranslationText to Speech

STEB: A Speech-to-Speech Translation Expressiveness Benchmark for Evaluating Beyond Translation Fidelity

2026-06-24 · Sitong Cheng, Weizhen Bian, Songjun Cao, Jin Li 외 arxiv

Speech-to-speech translation (S2ST) should preserve not only lexical meaning, but also expressive attributes: emotion, scenario style (e.g., news reporting vs. dramatic dialogue), and nonverbal vocalizations (NVs). Moreo…

Speech-to-Speech Translation