paper-with-me

홈 › Papers

Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation

2023-09-28 · Guy Yariv, Itai Gat, Sagie Benaim, Lior Wolf, Idan Schwartz, Yossi Adi

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio: globally, the input audio is semantically associated with the entire output video, and temporally, each segment of the input audio is associated with a corresponding segment of that video. We utilize an existing text-conditioned video generation model and a pre-trained audio encoder model. The proposed method is based on a lightweight adaptor network, which learns to map the audio-based representation to the input representation expected by the text-to-video generation model. As such, it also enables video generation conditioned on text, audio, and, for the first time as far as we can ascertain, on both text and audio. We validate our method extensively on three datasets demonstrating significant semantic diversity of audio-video samples and further propose a novel evaluation metric (AV-Align) to assess the alignment of generated videos with input audio samples. AV-Align is based on the detection and comparison of energy peaks in both modalities. In comparison to recent state-of-the-art approaches, our method generates videos that are better aligned with the input sound, both with respect to content and temporal axis. We also show that videos produced by our method present higher visual quality and are more diverse.

📄 PDF Abstract BibTeX arXiv:2309.16429

Code (1)

guyyariv/TempoTokens 공식 구현 pytorch

Tasks

Text-to-Video GenerationVideo Generation

Similar Papers 제목 키워드 기반

Audio-Sync Video Generation with Multi-Stream Temporal Control

2025-06-09 · Shuchen Weng, Haojie Zheng, Zheng Chang, Si Li 외

Audio is inherently temporal and closely synchronized with the visual world, making it a naturally aligned and expressive control signal for controllable video generation (e.g., movies). Beyond control, directly translat…

Audio-Visual SynchronizationVideo AlignmentVideo Generation

StereoSync: Spatially-Aware Stereo Audio Generation from Video

2025-10-07 · Christian Marinoni, Riccardo Fosco Gramaccioni, Kazuki Shimada, Takashi Shibuya 외 arxiv

Although audio generation has been widely studied over recent years, video-aligned audio generation still remains a relatively unexplored frontier. To address this gap, we introduce StereoSync, a novel and efficient mode…

Audio Generation

V2Meow: Meowing to the Visual Beat via Video-to-Music Generation

2023-05-11 · Kun Su, Judith Yue Li, Qingqing Huang, Dima Kuzmin 외

Video-to-music generation demands both a temporally localized high-quality listening experience and globally aligned video-acoustic signatures. While recent music generation models excel at the former through advanced au…

Music Generation

Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers

2025-09-26 · Jibin Song, Mingi Kwon, Jaeseok Jeong, Youngjung Uh arxiv

Text-to-video and image-to-video generation have made rapid progress in visual quality, but they remain limited in controlling the precise timing of motion. In contrast, audio provides temporal cues aligned with video mo…

Video Generation

Text-to-Audio Generation Synchronized with Videos

2024-03-08 · Shentong Mo, Jing Shi, Yapeng Tian

In recent times, the focus on text-to-audio (TTA) generation has intensified, as researchers strive to synthesize audio from textual descriptions. However, most existing methods, though leveraging latent diffusion models…

AudioCapsAudio GenerationContrastive Learning