paper-with-me

홈 › Papers

CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech

2025-06-03 · Helin Wang, Jiarui Hai, Dading Chong, Karan Thakkar, Tiantian Feng, Dongchao Yang, Junhyeok Lee, Laureano Moro Velazquez, Jesus Villalba, Zengyi Qin, Shrikanth Narayanan, Mounya Elhiali, Najim Dehak

Recent advancements in generative artificial intelligence have significantly transformed the field of style-captioned text-to-speech synthesis (CapTTS). However, adapting CapTTS to real-world applications remains challenging due to the lack of standardized, comprehensive datasets and limited research on downstream tasks built upon CapTTS. To address these gaps, we introduce CapSpeech, a new benchmark designed for a series of CapTTS-related tasks, including style-captioned text-to-speech synthesis with sound events (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS), and text-to-speech synthesis for chat agent (AgentTTS). CapSpeech comprises over 10 million machine-annotated audio-caption pairs and nearly 0.36 million human-annotated audio-caption pairs. In addition, we introduce two new datasets collected and recorded by a professional voice actor and experienced audio engineers, specifically for the AgentTTS and CapTTS-SE tasks. Alongside the datasets, we conduct comprehensive experiments using both autoregressive and non-autoregressive models on CapSpeech. Our results demonstrate high-fidelity and highly intelligible speech synthesis across a diverse range of speaking styles. To the best of our knowledge, CapSpeech is the largest available dataset offering comprehensive annotations for CapTTS-related tasks. The experiments and findings further provide valuable insights into the challenges of developing CapTTS systems.

📄 PDF Abstract BibTeX arXiv:2506.02863

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Similar Papers 제목 키워드 기반

How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech

2026-06-18 · Nityanand Mathur, Hamees Sayed, Wasim Madha, Apoorv Singh 외 arxiv

Style-captioned text-to-speech systems use natural language to control voice characteristics, but how individual words influence acoustic output remains unclear. Understanding this is critical for diagnosing failure mode…

Text-to-Image Synthesis Based on Machine Generated Captions

2019-10-09 · Marco Menardi, Alex Falcon, Saida S. Mohamed, Lorenzo Seidenari 외

Text to Image Synthesis refers to the process of automatic generation of a photo-realistic image starting from a given text and is revolutionizing many real-world applications. In order to perform such process it is nece…

Image CaptioningImage Generation

VectorFusion: Text-to-SVG by Abstracting Pixel-Based Diffusion Models

2022-11-21 · CVPR 2023 1 · Ajay Jain, Amber Xie, Pieter Abbeel

Diffusion models have shown impressive results in text-to-image synthesis. Using massive datasets of captioned images, diffusion models learn to generate raster images of highly diverse objects and scenes. However, desig…

Image GenerationText to 3DVector Graphics

Paint it Black: Generating paintings from text descriptions

2023-02-17 · Mahnoor Shahid, Mark Koch, Niklas Schneider

Two distinct tasks - generating photorealistic pictures from given text prompts and transferring the style of a painting to a real image to make it appear as though it were done by an artist, have been addressed many tim…

Image GenerationStyle Transfer

Neural Strokes: Stylized Line Drawing of 3D Shapes

2021-10-08 · ICCV 2021 10 · Difan Liu, Matthew Fisher, Aaron Hertzmann, Evangelos Kalogerakis

This paper introduces a model for producing stylized line drawings from 3D shapes. The model takes a 3D shape and a viewpoint as input, and outputs a drawing with textured strokes, with variations in stroke thickness, de…