paper-with-me

홈 › Papers

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

2024-06-04 · Philip Anastassiou, Jiawei Chen, Jitong Chen, Yuanzhe Chen, Zhuo Chen, Ziyi Chen, Jian Cong, Lelai Deng, Chuang Ding, Lu Gao, Mingqing Gong, Peisong Huang, Qingqing Huang, Zhiying Huang, YuanYuan Huo, Dongya Jia, ChuMin Li, Feiya Li, Hui Li, Jiaxin Li, Xiaoyang Li, Xingxing Li, Lin Liu, Shouda Liu, Sichao Liu, Xudong Liu, Yuchen Liu, Zhengxi Liu, Lu Lu, Junjie Pan, Xin Wang, Yuping Wang, Yuxuan Wang, Zhen Wei, Jian Wu, Chao Yao, Yifeng Yang, YuanHao Yi, Junteng Zhang, Qidi Zhang, Shuo Zhang, Wenjie Zhang, Yang Zhang, Zilin Zhao, Dejian Zhong, Xiaobin Zhuang

We introduce Seed-TTS, a family of large-scale autoregressive text-to-speech (TTS) models capable of generating speech that is virtually indistinguishable from human speech. Seed-TTS serves as a foundation model for speech generation and excels in speech in-context learning, achieving performance in speaker similarity and naturalness that matches ground truth human speech in both objective and subjective evaluations. With fine-tuning, we achieve even higher subjective scores across these metrics. Seed-TTS offers superior controllability over various speech attributes such as emotion and is capable of generating highly expressive and diverse speech for speakers in the wild. Furthermore, we propose a self-distillation method for speech factorization, as well as a reinforcement learning approach to enhance model robustness, speaker similarity, and controllability. We additionally present a non-autoregressive (NAR) variant of the Seed-TTS model, named $\text{Seed-TTS}_\text{DiT}$, which utilizes a fully diffusion-based architecture. Unlike previous NAR-based TTS systems, $\text{Seed-TTS}_\text{DiT}$ does not depend on pre-estimated phoneme durations and performs speech generation through end-to-end processing. We demonstrate that this variant achieves comparable performance to the language model-based variant and showcase its effectiveness in speech editing. We encourage readers to listen to demos at \url{https://bytedancespeech.github.io/seedtts_tech_report}.

📄 PDF Abstract BibTeX arXiv:2406.02430

Code (2)

BytedanceSpeech/seed-tts-eval 공식 구현 pytorch
Plachtaa/seed-vc pytorch

Tasks

In-Context LearningLanguage Modellingtext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Zero-shot Voice Conversion with Diffusion Transformers

2024-11-15 · Songting Liu

Zero-shot voice conversion aims to transform a source speech utterance to match the timbre of a reference speech from an unseen speaker. Traditional approaches struggle with timbre leakage, insufficient timbre representa…

In-Context LearningVoice Conversion

Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice

2025-07-23 · Shanbo Cheng, Yu Bao, Zhichao Huang, Yu Lu 외 arxiv

Simultaneous Interpretation (SI) represents one of the most daunting frontiers in the translation industry, with product-level automatic systems long plagued by intractable challenges: subpar transcription and translatio…

Speech-to-Speech TranslationReinforcement Learning

Diffiner: A Versatile Diffusion-based Generative Refiner for Speech Enhancement

2022-10-27 · Ryosuke Sawata, Naoki Murata, Yuhta Takida, Toshimitsu Uesaka 외

Although deep neural network (DNN)-based speech enhancement (SE) methods outperform the previous non-DNN-based ones, they often degrade the perceptual quality of generated outputs. To tackle this problem, we introduce a …

DenoisingSpeech Enhancement

SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing

2024-05-07 · Yuying Ge, Sijie Zhao, Chen Li, Yixiao Ge 외

In this technical report, we introduce SEED-Data-Edit: a unique hybrid dataset for instruction-guided image editing, which aims to facilitate image manipulation using open-form language. SEED-Data-Edit is composed of thr…

Image ManipulationLanguage ModelingLanguage ModellingLarge Language Model+1

SpeechX: Neural Codec Language Model as a Versatile Speech Transformer

2023-08-14 · Xiaofei Wang, Manthan Thakker, Zhuo Chen, Naoyuki Kanda 외

Recent advancements in generative speech models based on audio-text prompts have enabled remarkable innovations like high-quality zero-shot text-to-speech. However, existing models still face limitations in handling dive…

Language ModelingLanguage ModellingMulti-Task LearningSpeech Enhancement+3