paper-with-me

홈 › Papers

Expressive TTS Driven by Natural Language Prompts Using Few Human Annotations

2023-11-02 · Hanglei Zhang, Yiwei Guo, Sen Liu, Xie Chen, Kai Yu

Expressive text-to-speech (TTS) aims to synthesize speeches with human-like tones, moods, or even artistic attributes. Recent advancements in expressive TTS empower users with the ability to directly control synthesis style through natural language prompts. However, these methods often require excessive training with a significant amount of style-annotated data, which can be challenging to acquire. Moreover, they may have limited adaptability due to fixed style annotations. In this work, we present FreeStyleTTS (FS-TTS), a controllable expressive TTS model with minimal human annotations. Our approach utilizes a large language model (LLM) to transform expressive TTS into a style retrieval task. The LLM selects the best-matching style references from annotated utterances based on external style prompts, which can be raw input text or natural language style descriptions. The selected reference guides the TTS pipeline to synthesize speeches with the intended style. This innovative approach provides flexible, versatile, and precise style control with minimal human workload. Experiments on a Mandarin storytelling corpus demonstrate FS-TTS's proficiency in leveraging LLM's semantic inference ability to retrieve desired styles from either input text or user-defined descriptions. This results in synthetic speeches that are closely aligned with the specified styles.

📄 PDF Abstract BibTeX arXiv:2311.01260

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelRetrievaltext-to-speechText to Speech

Similar Papers 제목 키워드 기반

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech

2025-05-20 · Yu Pan, Yanni Hu, Yuguang Yang, Jixun Yao 외

Despite great advances, achieving high-fidelity emotional voice conversion (EVC) with flexible and interpretable control remains challenging. This paper introduces ClapFM-EVC, a novel EVC framework capable of generating …

Voice Conversion

MOSS-VoiceGenerator: Create Realistic Voices with Natural Language Descriptions

2026-03-30 · Kexin Huang, Liwei Fan, Botian Jiang, Yaozhou Jiang 외 arxiv

Voice design from natural language aims to generate speaker timbres directly from free-form textual descriptions, allowing users to create voices tailored to specific roles, personalities, and emotions. Such controllable…

PromptVC: Flexible Stylistic Voice Conversion in Latent Space Driven by Natural Language Prompts

2023-09-17 · Jixun Yao, Yuguang Yang, Yi Lei, Ziqian Ning 외

Style voice conversion aims to transform the style of source speech to a desired style according to real-world application demands. However, the current style voice conversion approach relies on pre-defined labels or ref…

Voice Conversion

Action2Dialogue: Generating Character-Centric Narratives from Scene-Level Prompts

2025-05-22 · Taewon Kang, Ming C. Lin

Recent advances in scene-based video generation have enabled systems to synthesize coherent visual narratives from structured prompts. However, a crucial dimension of storytelling -- character-driven dialogue and speech …

Dialogue GenerationLarge Language ModelStory GenerationVideo Generation+1

SpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language Description

2024-08-24 · Zeyu Jin, Jia Jia, Qixin Wang, Kehan Li 외

Speech-language multi-modal learning presents a significant challenge due to the fine nuanced information inherent in speech styles. Therefore, a large-scale dataset providing elaborate comprehension of speech style is u…

DescriptiveSpeech SynthesisTAG