paper-with-me

홈 › Papers

Guiding Audio Editing with Audio Language Model

2025-09-25 · Zitong Lan, Yiduo Hao, Mingmin Zhao arxiv

Audio editing plays a central role in VR/AR immersion, virtual conferencing, sound design, and other interactive media. However, recent generative audio editing models depend on template-like instruction formats and are restricted to mono-channel audio. These models fail to deal with declarative audio editing, where the user declares what the desired outcome should be, while leaving the details of editing operations to the system. We introduce SmartDJ, a novel framework for stereo audio editing that combines the reasoning capability of audio language models with the generative power of latent diffusion. Given a high-level instruction, SmartDJ decomposes it into a sequence of atomic edit operations, such as adding, removing, or spatially relocating events. These operations are then executed by a diffusion model trained to manipulate stereo audio. To support this, we design a data synthesis pipeline that produces paired examples of high-level instructions, atomic edit operations, and audios before and after each edit operation. Experiments demonstrate that SmartDJ achieves superior perceptual quality, spatial realism, and semantic alignment compared to prior audio editing methods. Demos are available at https://zitonglan.github.io/project/smartdj/smartdj.html.

📄 PDF Abstract BibTeX arXiv:2509.21625

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SAO-Instruct: Free-form Audio Editing using Natural Language Instructions

2025-10-26 · Michael Ungersböck, Florian Grötschla, Luca A. Lanzendörfer, June Young Yi 외 arxiv

Generative models have made significant progress in synthesizing high-fidelity audio from short textual descriptions. However, editing existing audio using natural language has remained largely underexplored. Current app…

Language-Guided Joint Audio-Visual Editing via One-Shot Adaptation

2024-10-09 · Susan Liang, Chao Huang, Yapeng Tian, Anurag Kumar 외

In this paper, we introduce a novel task called language-guided joint audio-visual editing. Given an audio and image pair of a sounding event, this task aims at generating new audio-visual content by editing the given so…

DirectAudioEdit: Inversion-Free Text-Guided Audio Editing via Diffusion Prediction Contrast

2026-06-05 · Zhengkun Ge, Xiaoqian Liu, Haoran Zhang, Yuan Ge 외 arxiv

Text-guided audio editing aims to modify the language-specified acoustic content while preserving edit-irrelevant source components. Existing training-free methods typically rely on inversion-based editing. While inversi…

Instruct-NeuralTalker: Editing Audio-Driven Talking Radiance Fields with Instructions

2023-06-19 · Yuqi Sun, Ruian He, Weimin Tan, Bo Yan

Recent neural talking radiance field methods have shown great success in photorealistic audio-driven talking face synthesis. In this paper, we propose a novel interactive framework that utilizes human instructions to edi…

Face GenerationTalking Face Generation

Localizing and Editing Knowledge in Large Audio-Language Models

2026-03-15 · Sung Kyun Chung, Jiaheng Dong, Qiuchi Hu, Gongping Huang 외 arxiv

Large Audio-Language Models (LALMs) have shown strong performance in speech understanding, making speech a natural interface for accessing factual information. Yet they are trained on static corpora and may encode incorr…