paper-with-me

Papers

Virtual Consistency for Audio Editing

2025-09-21 · Matthieu Cervera, Francesco Paissan, Mirco Ravanelli, Cem Subakan arxiv

Free-form, text-based audio editing remains a persistent challenge, despite progress in inversion-based neural methods. Current approaches rely on slow inversion procedures, limiting their practicality. We present a virtual-consistency based audio editing system that bypasses inversion by adapting the sampling process of diffusion models. Our pipeline is model-agnostic, requiring no fine-tuning or architectural changes, and achieves substantial speed-ups over recent neural editing baselines. Crucially, it achieves this efficiency without compromising quality, as demonstrated by quantitative benchmarks and a user study involving 16 participants.

📄 PDF Abstract BibTeX arXiv:2509.17219

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AudioScenic: Audio-Driven Video Scene Editing

2024-04-25 · Kaixin Shen, Ruijie Quan, Linchao Zhu, Jun Xiao 외

Audio-driven visual scene editing endeavors to manipulate the visual background while leaving the foreground content unchanged, according to the given audio signals. Unlike current efforts focusing primarily on image edi…

Guiding Audio Editing with Audio Language Model

2025-09-25 · Zitong Lan, Yiduo Hao, Mingmin Zhao arxiv

Audio editing plays a central role in VR/AR immersion, virtual conferencing, sound design, and other interactive media. However, recent generative audio editing models depend on template-like instruction formats and are …

OmniEdit: A Training-free framework for Lip Synchronization and Audio-Visual Editing

2026-03-10 · Lixiang Lin, Siyuan Jin, Jinshan Zhang arxiv

Lip synchronization and audio-visual editing have emerged as fundamental challenges in multimodal learning, underpinning a wide range of applications, including film production, virtual avatars, and telepresence. Despite…

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

2026-07-17 · Yuqing Wen, Yukai Huang, Qianqian Xie, Jiangtao Wu 외 hf

While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmark…

Instruction Following

Language-Guided Joint Audio-Visual Editing via One-Shot Adaptation

2024-10-09 · Susan Liang, Chao Huang, Yapeng Tian, Anurag Kumar 외

In this paper, we introduce a novel task called language-guided joint audio-visual editing. Given an audio and image pair of a sounding event, this task aims at generating new audio-visual content by editing the given so…