paper-with-me

홈 › Papers

Coherent Audio-Visual Editing via Conditional Audio Generation Following Video Edits

2025-12-08 · Masato Ishii, Akio Hayakawa, Takashi Shibuya, Yuki Mitsufuji arxiv

We introduce a novel pipeline for joint audio-visual editing that enhances the coherence between edited video and its accompanying audio. Our approach first applies state-of-the-art video editing techniques to produce the target video, then performs audio editing to align with the visual changes. To achieve this, we present a new video-to-audio generation model that conditions on the source audio, target video, and a text prompt. We extend the model architecture to incorporate conditional audio input and propose a data augmentation strategy that improves training efficiency. Furthermore, our model dynamically adjusts the influence of the source audio based on the complexity of the edits, preserving the original audio structure where possible. Experimental results demonstrate that our method outperforms existing approaches in maintaining audio-visual alignment and content integrity.

📄 PDF Abstract BibTeX arXiv:2512.07209

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationAudio Generation

Similar Papers 제목 키워드 기반

Text-based Talking Video Editing with Cascaded Conditional Diffusion

2024-07-20 · Bo Han, Heqing Zou, Haoyang Li, Guangcong Wang 외

Text-based talking-head video editing aims to efficiently insert, delete, and substitute segments of talking videos through a user-friendly text editing approach. It is challenging because of \textbf{1)} generalizable ta…

Video Editing

Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising

2025-03-26 · Yan-Bo Lin, Kevin Lin, Zhengyuan Yang, Linjie Li 외

In this paper, we introduce zero-shot audio-video editing, a novel task that requires transforming original audio-visual content to align with a specified textual prompt without additional model training. To evaluate thi…

DenoisingVideo Editing

ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing

2025-06-26 · Huadai Liu, Jialei Wang, Kaicheng Luo, Wen Wang 외

While end-to-end video-to-audio generation has greatly improved, producing high-fidelity audio that authentically captures the nuances of visual content remains challenging. Like professionals in the creative industries,…

Audio GenerationLarge Language ModelMultimodal Large Language Model

Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows

2026-03-09 · Shentong Mo, Yibing Song arxiv

Coordinated audio generation based on video inputs typically requires a strict audio-visual (AV) alignment, where both semantics and rhythmics of the generated audio segments shall correspond to those in the video frames…

Contrastive LearningAudio Generation

Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows

2025-01-01 · CVPR 2025 1 · Shentong Mo, Yibing Song

Coordinated audio generation based on video inputs typically requires a strict audio-visual (AV) alignment, where both semantics and rhythmics of the generated audio segments shall correspond to those in the video fr…

Audio GenerationContrastive Learning