paper-with-me

홈 › Papers

AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions

2024-09-19 · Yuanyuan Wang, Hangting Chen, Dongchao Yang, Zhiyong Wu, Xixin Wu

Current Text-to-audio (TTA) models mainly use coarse text descriptions as inputs to generate audio, which hinders models from generating audio with fine-grained control of content and style. Some studies try to improve the granularity by incorporating additional frame-level conditions or control networks. However, this usually leads to complex system design and difficulties due to the requirement for reference frame-level conditions. To address these challenges, we propose AudioComposer, a novel TTA generation framework that relies solely on natural language descriptions (NLDs) to provide both content specification and style control information. To further enhance audio generative modeling, we employ flow-based diffusion transformers with the cross-attention mechanism to incorporate text descriptions effectively into audio generation processes, which can not only simultaneously consider the content and style information in the text inputs, but also accelerate generation compared to other architectures. Furthermore, we propose a novel and comprehensive automatic data simulation pipeline to construct data with fine-grained text descriptions, which significantly alleviates the problem of data scarcity in the area. Experiments demonstrate the effectiveness of our framework using solely NLDs as inputs for content specification and style control. The generation quality and controllability surpass state-of-the-art TTA models, even with a smaller model size.

📄 PDF Abstract BibTeX arXiv:2409.12560

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

2026-08-03 · Yu Zhang, Ruiqi Li, Changhao Pan, Ke Lei 외 hf

Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference re…

Audio Generation

OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation

2025-06-23 · Qijun Gan, Ruizi Yang, Jianke Zhu, Shaofei Xue 외

Significant progress has been made in audio-driven human animation, while most existing methods focus mainly on facial movements, limiting their ability to create full-body animations with natural synchronization and flu…

Human AnimationVideo Generation

SpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language Description

2024-08-24 · Zeyu Jin, Jia Jia, Qixin Wang, Kehan Li 외

Speech-language multi-modal learning presents a significant challenge due to the fine nuanced information inherent in speech styles. Therefore, a large-scale dataset providing elaborate comprehension of speech style is u…

DescriptiveSpeech SynthesisTAG

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning

2026-07-22 · Siqian Tong, Xuan Li, Chaozhuo Li, Baolong Bi 외 arxiv

Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-t…

AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer

2026-03-16 · Pengjun Fang, Yingqing He, Yazhou Xing, Qifeng Chen 외 arxiv

Existing video-to-audio (V2A) generation methods predominantly rely on text prompts alongside visual information to synthesize audio. However, two critical bottlenecks persist: semantic granularity gaps in training data,…