paper-with-me

Papers

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts

2025-01-08 · Hanzhao Li, Yuke Li, Xinsheng Wang, Jingbin Hu, Qicong Xie, Shan Yang, Lei Xie

Controllable speech generation methods typically rely on single or fixed prompts, hindering creativity and flexibility. These limitations make it difficult to meet specific user needs in certain scenarios, such as adjusting the style while preserving a selected speaker's timbre, or choosing a style and generating a voice that matches a character's visual appearance. To overcome these challenges, we propose \textit{FleSpeech}, a novel multi-stage speech generation framework that allows for more flexible manipulation of speech attributes by integrating various forms of control. FleSpeech employs a multimodal prompt encoder that processes and unifies different text, audio, and visual prompts into a cohesive representation. This approach enhances the adaptability of speech synthesis and supports creative and precise control over the generated speech. Additionally, we develop a data collection pipeline for multimodal datasets to facilitate further research and applications in this field. Comprehensive subjective and objective experiments demonstrate the effectiveness of FleSpeech. Audio samples are available at https://kkksuper.github.io/FleSpeech/

📄 PDF Abstract BibTeX arXiv:2501.04644

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

FlexEControl: Flexible and Efficient Multimodal Control for Text-to-Image Generation

2024-05-08 · Xuehai He, Jian Zheng, Jacob Zhiyuan Fang, Robinson Piramuthu 외

Controllable text-to-image (T2I) diffusion models generate images conditioned on both text prompts and semantic inputs of other modalities like edge maps. Nevertheless, current controllable T2I methods commonly face chal…

Image GenerationText to Image GenerationText-to-Image Generation

FlexCAD: Unified and Versatile Controllable CAD Generation with Fine-tuned Large Language Models

2024-11-05 · Zhanwei Zhang, Shizhao Sun, Wenxiao Wang, Deng Cai 외

Recently, there is a growing interest in creating computer-aided design (CAD) models based on user intent, known as controllable CAD generation. Existing work offers limited controllability and needs separate models for …

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Survey

2024-12-09 · Tianxin Xie, Yan Rong, Pengfei Zhang, Wenwu Wang 외

Text-to-speech (TTS), also known as speech synthesis, is a prominent research area that aims to generate natural-sounding human speech from text. Recently, with the increasing industrial demand, TTS technologies have evo…

Speech SynthesisSurveytext-to-speechText to Speech

Controllable Data Generation by Deep Learning: A Review

2022-07-19 · Shiyu Wang, Yuanqi Du, Xiaojie Guo, Bo Pan 외

Designing and generating new data under targeted properties has been attracting various critical applications such as molecule design, image editing and speech synthesis. Traditional hand-crafted approaches heavily rely …

Deep LearningSpeech Synthesis

FilterPrompt: A Simple yet Efficient Approach to Guide Image Appearance Transfer in Diffusion Models

2024-04-20 · Xi Wang, Yichen Peng, Heng Fang, Yilin Wang 외

In controllable generation tasks, flexibly manipulating the generated images to attain a desired appearance or structure based on a single input image cue remains a critical and longstanding challenge. Achieving this req…

Appearance TransferFeature Correlation