paper-with-me

홈 › Papers

On The Open Prompt Challenge In Conditional Audio Generation

2023-11-01 · Ernie Chang, Sidd Srinivasan, Mahi Luthra, Pin-Jie Lin, Varun Nagaraja, Forrest Iandola, Zechun Liu, Zhaoheng Ni, Changsheng Zhao, Yangyang Shi, Vikas Chandra

Text-to-audio generation (TTA) produces audio from a text description, learning from pairs of audio samples and hand-annotated text. However, commercializing audio generation is challenging as user-input prompts are often under-specified when compared to text descriptions used to train TTA models. In this work, we treat TTA models as a `blackbox'' and address the user prompt challenge with two key insights: (1) User prompts are generally under-specified, leading to a large alignment gap between user prompts and training prompts. (2) There is a distribution of audio descriptions for which TTA models are better at generating higher quality audio, which we refer to as `audionese''. To this end, we rewrite prompts with instruction-tuned models and propose utilizing text-audio alignment as feedback signals via margin ranking learning for audio improvements. On both objective and subjective human evaluations, we observed marked improvements in both text-audio alignment and music audio quality.

📄 PDF Abstract BibTeX arXiv:2311.00897

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Generation

Similar Papers 제목 키워드 기반

In-Context Prompt Editing For Conditional Audio Generation

2023-11-01 · Ernie Chang, Pin-Jie Lin, Yang Li, Sidd Srinivasan 외

Distributional shift is a central challenge in the deployment of machine learning models as they can be ill-equipped for real-world data. This is particularly evident in text-to-audio generation where the encoded represe…

Audio GenerationRetrieval

OpenDance: Multimodal Controllable 3D Dance Generation Using Large-scale Internet Data

2025-06-09 · Jinlu Zhang, Zixi Kang, Yizhou Wang

Music-driven dance generation offers significant creative potential yet faces considerable challenges. The absence of fine-grained multimodal data and the difficulty of flexible multi-conditional generation limit previou…

Diversity

PTQ4ADM: Post-Training Quantization for Efficient Text Conditional Audio Diffusion Models

2024-09-20 · Jayneel Vora, Aditya Krishnan, Nader Bouacida, Prabhu RV Shankar 외

Denoising diffusion models have emerged as state-of-the-art in generative tasks across image, audio, and video domains, producing high-quality, diverse, and contextually relevant data. However, their broader adoption is …

Audio GenerationAudio SynthesisDenoisingQuantization

Efficient Parallel Audio Generation using Group Masked Language Modeling

2024-01-02 · Myeonghun Jeong, Minchan Kim, Joun Yeop Lee, Nam Soo Kim

We present a fast and high-quality codec language model for parallel audio generation. While SoundStorm, a state-of-the-art parallel audio generation model, accelerates inference speed compared to autoregressive models, …

Audio GenerationComputational EfficiencyLanguage ModelingLanguage Modelling+1

Coherent Audio-Visual Editing via Conditional Audio Generation Following Video Edits

2025-12-08 · Masato Ishii, Akio Hayakawa, Takashi Shibuya, Yuki Mitsufuji arxiv

We introduce a novel pipeline for joint audio-visual editing that enhances the coherence between edited video and its accompanying audio. Our approach first applies state-of-the-art video editing techniques to produce th…

Data AugmentationAudio Generation