paper-with-me

Papers

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance

2025-06-26 · Akio Hayakawa, Masato Ishii, Takashi Shibuya, Yuki Mitsufuji

We propose a novel step-by-step video-to-audio generation method that sequentially produces individual audio tracks, each corresponding to a specific sound event in the video. Our approach mirrors traditional Foley workflows, aiming to capture all sound events induced by a given video comprehensively. Each generation step is formulated as a guided video-to-audio synthesis task, conditioned on a target text prompt and previously generated audio tracks. This design is inspired by the idea of concept negation from prior compositional generation frameworks. To enable this guided generation, we introduce a training framework that leverages pre-trained video-to-audio models and eliminates the need for specialized paired datasets, allowing training on more accessible data. Experimental results demonstrate that our method generates multiple semantically distinct audio tracks for a single input video, leading to higher-quality composite audio synthesis than existing baselines.

📄 PDF Abstract BibTeX arXiv:2506.20995

Code (0)

등록된 구현이 없습니다.

Tasks

Audio GenerationAudio SynthesisNegation

Similar Papers 제목 키워드 기반

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation

2025-09-08 · Xiaoran Yang, Jianxuan Yang, Xinyue Guo, Haoyu Wang 외 arxiv

A key challenge in synthesizing audios from silent videos is the inherent trade-off between synthesis quality and inference efficiency in existing methods. For instance, flow matching based models rely on modeling instan…

Large-scale unsupervised audio pre-training for video-to-speech synthesis

2023-06-27 · Triantafyllos Kefalas, Yannis Panagakis, Maja Pantic

Video-to-speech synthesis is the task of reconstructing the speech signal from a silent video of a speaker. Most established approaches to date involve a two-step process, whereby an intermediate representation from the …

speech-recognitionSpeech RecognitionSpeech Synthesis

TARO: Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning for Synchronized Video-to-Audio Synthesis

2025-04-08 · Tri Ton, Ji Woo Hong, Chang D. Yoo

This paper introduces Timestep-Adaptive Representation Alignment with Onset-Aware Conditioning (TARO), a novel framework for high-fidelity and temporally coherent video-to-audio synthesis. Built upon flow-based transform…

Audio SynthesisFAD

Long-Video Audio Synthesis with Multi-Agent Collaboration

2025-03-13 · Yehang Zhang, Xinli Xu, Xiaojie Xu, Li Liu 외

Video-to-audio synthesis, which generates synchronized audio for visual content, critically enhances viewer immersion and narrative coherence in film and interactive media. However, video-to-audio dubbing for long-form c…

Audio SynthesisScene SegmentationScript Generation

Audeo: Audio Generation for a Silent Performance Video

2020-06-23 · NeurIPS 2020 12 · Kun Su, Xiulong Liu, Eli Shlizerman

We present a novel system that gets as an input video frames of a musician playing the piano and generates the music for that video. Generation of music from visual cues is a challenging problem and it is not clear wheth…

Audio GenerationAudio SynthesisVideo Generation