paper-with-me

홈 › Papers

AUDIT: Audio Editing by Following Instructions with Latent Diffusion Models

2023-04-03 · NeurIPS 2023 11

Audio editing is applicable for various purposes, such as adding background sound effects, replacing a musical instrument, and repairing damaged audio. Recently, some diffusion-based methods achieved zero-shot audio editing by using a diffusion and denoising process conditioned on the text description of the output audio. However, these methods still have some problems: 1) they have not been trained on editing tasks and cannot ensure good editing effects; 2) they can erroneously modify audio segments that do not require editing; 3) they need a complete description of the output audio, which is not always available or necessary in practical scenarios. In this work, we propose AUDIT, an instruction-guided audio editing model based on latent diffusion models. Specifically, AUDIT has three main design features: 1) we construct triplet training data (instruction, input audio, output audio) for different audio editing tasks and train a diffusion model using instruction and input (to be edited) audio as conditions and generating output (edited) audio; 2) it can automatically learn to only modify segments that need to be edited by comparing the difference between the input and output audio; 3) it only needs edit instructions instead of full target audio descriptions as text input. AUDIT achieves state-of-the-art results in both objective and subjective metrics for several audio editing tasks (e.g., adding, dropping, replacement, inpainting, super-resolution). Demo samples are available at https://audit-demo.github.io/.

📄 PDF Abstract BibTeX arXiv:2304.00830

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingSuper-ResolutionTriplet

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

InstructAny2Pix: Flexible Visual Editing via Multimodal Instruction Following

2023-12-11 · Shufan Li, Harkanwar Singh, Aditya Grover

The ability to provide fine-grained control for generating and editing visual imagery has profound implications for computer vision and its applications. Previous works have explored extending controllability in two dire…

DecoderInstruction Following

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

2026-07-17 · Yuqing Wen, Yukai Huang, Qianqian Xie, Jiangtao Wu 외 hf

While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmark…

Instruction Following

Audio Texture Manipulation by Exemplar-Based Analogy

2025-01-21 · Kan Jen Cheng, Tingle Li, Gopala Anumanchipalli

Audio texture manipulation involves modifying the perceptual characteristics of a sound to achieve specific transformations, such as adding, removing, or replacing auditory elements. In this paper, we propose an exemplar…

AudioMorphix: Training-free audio editing with diffusion probabilistic models

2025-05-21 · Jinhua Liang, Yuanzhe Chen, Yi Yuan, Dongya Jia 외

Editing sound with precision is a crucial yet underexplored challenge in audio content creation. While existing works can manipulate sounds by text instructions or audio exemplar pairs, they often struggled to modify aud…

Guiding Audio Editing with Audio Language Model

2025-09-25 · Zitong Lan, Yiduo Hao, Mingmin Zhao arxiv

Audio editing plays a central role in VR/AR immersion, virtual conferencing, sound design, and other interactive media. However, recent generative audio editing models depend on template-like instruction formats and are …