paper-with-me

홈 › Papers

Temporally Aligned Audio for Video with Autoregression

2024-09-20 · Ilpo Viertola, Vladimir Iashin, Esa Rahtu

We introduce V-AURA, the first autoregressive model to achieve high temporal alignment and relevance in video-to-audio generation. V-AURA uses a high-framerate visual feature extractor and a cross-modal audio-visual feature fusion strategy to capture fine-grained visual motion events and ensure precise temporal alignment. Additionally, we propose VisualSound, a benchmark dataset with high audio-visual relevance. VisualSound is based on VGGSound, a video dataset consisting of in-the-wild samples extracted from YouTube. During the curation, we remove samples where auditory events are not aligned with the visual ones. V-AURA outperforms current state-of-the-art models in temporal alignment and semantic relevance while maintaining comparable audio quality. Code, samples, VisualSound and models are available at https://v-aura.notion.site

📄 PDF Abstract BibTeX arXiv:2409.13689

Code (1)

ilpoviertola/V-AURA 공식 구현 pytorch

Tasks

Audio GenerationVideo-to-Sound Generation

Similar Papers 제목 키워드 기반

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation

2024-12-19 · Moayed Haji-Ali, Willi Menapace, Aliaksandr Siarohin, Ivan Skorokhodov 외

We propose AV-Link, a unified framework for Video-to-Audio (A2V) and Audio-to-Video (A2V) generation that leverages the activations of frozen video and audio diffusion models for temporally-aligned cross-modal conditioni…

Video GenerationVideo Synchronization

Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation

2023-09-28 · Guy Yariv, Itai Gat, Sagie Benaim, Lior Wolf 외

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally w…

Text-to-Video GenerationVideo Generation

Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows

2026-03-09 · Shentong Mo, Yibing Song arxiv

Coordinated audio generation based on video inputs typically requires a strict audio-visual (AV) alignment, where both semantics and rhythmics of the generated audio segments shall correspond to those in the video frames…

Contrastive LearningAudio Generation

Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows

2025-01-01 · CVPR 2025 1 · Shentong Mo, Yibing Song

Coordinated audio generation based on video inputs typically requires a strict audio-visual (AV) alignment, where both semantics and rhythmics of the generated audio segments shall correspond to those in the video fr…

Audio GenerationContrastive Learning

Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models

2023-06-29 · NeurIPS 2023 11 · Simian Luo, Chuanhao Yan, Chenxu Hu, Hang Zhao

The Video-to-Audio (V2A) model has recently gained attention for its practical application in generating audio directly from silent videos, particularly in video/film production. However, previous methods in V2A have lim…

Audio Synthesis