paper-with-me

홈 › Papers

Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis

2023-07-14 · Ziyue Jiang, Jinglin Liu, Yi Ren, Jinzheng He, Zhenhui Ye, Shengpeng Ji, Qian Yang, Chen Zhang, Pengfei Wei, Chunfeng Wang, Xiang Yin, Zejun Ma, Zhou Zhao

Zero-shot text-to-speech (TTS) aims to synthesize voices with unseen speech prompts, which significantly reduces the data and computation requirements for voice cloning by skipping the fine-tuning process. However, the prompting mechanisms of zero-shot TTS still face challenges in the following aspects: 1) previous works of zero-shot TTS are typically trained with single-sentence prompts, which significantly restricts their performance when the data is relatively sufficient during the inference stage. 2) The prosodic information in prompts is highly coupled with timbre, making it untransferable to each other. This paper introduces Mega-TTS 2, a generic prompting mechanism for zero-shot TTS, to tackle the aforementioned challenges. Specifically, we design a powerful acoustic autoencoder that separately encodes the prosody and timbre information into the compressed latent space while providing high-quality reconstructions. Then, we propose a multi-reference timbre encoder and a prosody latent language model (P-LLM) to extract useful information from multi-sentence prompts. We further leverage the probabilities derived from multiple P-LLM outputs to produce transferable and controllable prosody. Experimental results demonstrate that Mega-TTS 2 could not only synthesize identity-preserving speech with a short prompt of an unseen speaker from arbitrary sources but consistently outperform the fine-tuning method when the volume of data ranges from 10 seconds to 5 minutes. Furthermore, our method enables to transfer various speaking styles to the target timbre in a fine-grained and controlled manner. Audio samples can be found in https://boostprompt.github.io/boostprompt/.

📄 PDF Abstract BibTeX arXiv:2307.07218

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context LearningLanguage ModellingSentenceSpeech Synthesistext-to-speechText to SpeechVoice Cloning

Similar Papers 제목 키워드 기반

LSTPrompt: Large Language Models as Zero-Shot Time Series Forecasters by Long-Short-Term Prompting

2024-02-25 · Haoxin Liu, Zhiyuan Zhao, Jindong Wang, Harshavardhan Kamarthi 외

Time-series forecasting (TSF) finds broad applications in real-world scenarios. Prompting off-the-shelf Large Language Models (LLMs) demonstrates strong zero-shot TSF capabilities while preserving computational efficienc…

Computational EfficiencyTime SeriesTime Series Forecasting

MegaFlow: Zero-Shot Large Displacement Optical Flow

2026-03-26 · Dingxi Zhang, Fangjinhua Wang, Marc Pollefeys, Haofei Xu arxiv

Accurate estimation of large displacement optical flow remains a critical challenge. Existing methods typically rely on iterative local search or/and domain-specific fine-tuning, which severely limits their performance i…

Zero-shot GeneralizationPoint Tracking

Do Prompts Reshape Representations? An Empirical Study of Prompting Effects on Embeddings

2025-10-22 · Cesar Gonzalez-Gutierrez, Dirk Hovy arxiv

Prompting is a common approach for leveraging LMs in zero-shot settings. However, the underlying mechanisms that enable LMs to perform diverse tasks without task-specific supervision remain poorly understood. Studying th…

Dynamic Strategy Chain: Dynamic Zero-Shot CoT for Long Mental Health Support Generation

2023-08-21 · Qi Chen, Dexi Liu

Long counseling Text Generation for Mental health support (LTGM), an innovative and challenging task, aims to provide help-seekers with mental health support through a comprehensive and more acceptable response. The comb…

Text Generation

Boosting Zero-Shot Human-Object Interaction Detection with Vision-Language Transfer

2024-03-18 · IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2024 3 · Sandipan Sarma, Pradnesh Kalkar, Arijit Sur

Human-Object Interaction (HOI) detection is a crucial task that involves localizing interactive human-object pairs and identifying the actions being performed. Most existing HOI detectors are supervised in nature and lac…

Human-Object Interaction DetectionLanguage ModelingLanguage ModellingObject+1