paper-with-me

홈 › Papers

In-Context Prompt Editing For Conditional Audio Generation

2023-11-01 · Ernie Chang, Pin-Jie Lin, Yang Li, Sidd Srinivasan, Gael Le Lan, David Kant, Yangyang Shi, Forrest Iandola, Vikas Chandra

Distributional shift is a central challenge in the deployment of machine learning models as they can be ill-equipped for real-world data. This is particularly evident in text-to-audio generation where the encoded representations are easily undermined by unseen prompts, which leads to the degradation of generated audio -- the limited set of the text-audio pairs remains inadequate for conditional audio generation in the wild as user prompts are under-specified. In particular, we observe a consistent audio quality degradation in generated audio samples with user prompts, as opposed to training set prompts. To this end, we present a retrieval-based in-context prompt editing framework that leverages the training captions as demonstrative exemplars to revisit the user prompts. We show that the framework enhanced the audio quality across the set of collected user prompts, which were edited with reference to the training captions as exemplars.

📄 PDF Abstract BibTeX arXiv:2311.00895

Code (0)

등록된 구현이 없습니다.

Tasks

Audio GenerationRetrieval

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Coherent Audio-Visual Editing via Conditional Audio Generation Following Video Edits

2025-12-08 · Masato Ishii, Akio Hayakawa, Takashi Shibuya, Yuki Mitsufuji arxiv

We introduce a novel pipeline for joint audio-visual editing that enhances the coherence between edited video and its accompanying audio. Our approach first applies state-of-the-art video editing techniques to produce th…

Data AugmentationAudio Generation

Text-based Talking Video Editing with Cascaded Conditional Diffusion

2024-07-20 · Bo Han, Heqing Zou, Haoyang Li, Guangcong Wang 외

Text-based talking-head video editing aims to efficiently insert, delete, and substitute segments of talking videos through a user-friendly text editing approach. It is challenging because of \textbf{1)} generalizable ta…

Video Editing

Audio Editing with Non-Rigid Text Prompts

2023-10-19 · Francesco Paissan, Luca Della Libera, Zhepei Wang, Mirco Ravanelli 외

In this paper, we explore audio-editing with non-rigid text edits. We show that the proposed editing pipeline is able to create audio edits that remain faithful to the input audio. We explore text prompts that perform ad…

Audio GenerationStyle Transfer

Prompt-guided Precise Audio Editing with Diffusion Models

2024-05-11 · Manjie Xu, Chenxing Li, Duzhen Zhang, Dan Su 외

Audio editing involves the arbitrary manipulation of audio content through precise control. Although text-guided diffusion models have made significant advancements in text-to-audio generation, they still face challenges…

Audio Generation

CAMEO: A Conditional and Quality-Aware Multi-Agent Image Editing Orchestrator

2026-04-03 · Yuhan Pu, Hao Zheng, Ziqian Mo, Zirui Pang 외 arxiv

Conditional image editing aims to modify a source image according to textual prompts and optional reference guidance. Such editing is crucial in scenarios requiring strict structural control (i.e., anomaly insertion in d…

Image Editing