paper-with-me

Papers

Sound-Guided Semantic Video Generation

2022-04-20 · Seung Hyun Lee, Gyeongrok Oh, Wonmin Byeon, Chanyoung Kim, Won Jeong Ryoo, Sang Ho Yoon, Hyunjun Cho, Jihyun Bae, Jinkyu Kim, Sangpil Kim

The recent success in StyleGAN demonstrates that pre-trained StyleGAN latent space is useful for realistic video generation. However, the generated motion in the video is usually not semantically meaningful due to the difficulty of determining the direction and magnitude in the StyleGAN latent space. In this paper, we propose a framework to generate realistic videos by leveraging multimodal (sound-image-text) embedding space. As sound provides the temporal contexts of the scene, our framework learns to generate a video that is semantically consistent with sound. First, our sound inversion module maps the audio directly into the StyleGAN latent space. We then incorporate the CLIP-based multimodal embedding space to further provide the audio-visual relationships. Finally, the proposed frame generator learns to find the trajectory in the latent space which is coherent with the corresponding sound and generates a video in a hierarchical manner. We provide the new high-resolution landscape video dataset (audio-visual pair) for the sound-guided video generation task. The experiments show that our model outperforms the state-of-the-art methods in terms of video quality. We further show several applications including image and video editing to verify the effectiveness of our method.

📄 PDF Abstract BibTeX arXiv:2204.09273

Code (0)

등록된 구현이 없습니다.

Tasks

Video EditingVideo Generation

Methods 이 논문이 사용한 방법론

StyleGAN 설명 없음
Adaptive Instance Normalization 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
HuMan(Expedia)||How do I get a human at Expedia? How do I get a human at Expedia? How Do I Get a Human at Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Real-Time Help & Exclusive…
R1 Regularization R_INLINE_MATH_1 Regularization is a regularization technique and gradient penalty for training [generative adversarial…
Feedforward Network A Feedforward Network, or a Multilayer Perceptron (MLP), is a neural network with solely densely connected layers. This is the classic neural network architecture of the…

Similar Papers 제목 키워드 기반

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls

2024-12-12 · Zihao Chen, Haomin Zhang, Xinhan Di, Haoyu Wang 외

Generating sound effects for product-level videos, where only a small amount of labeled data is available for diverse scenes, requires the production of high-quality sounds in few-shot settings. To tackle the challenge o…

Audio Generation

The Power of Sound (TPoS): Audio Reactive Video Generation with Stable Diffusion

2023-09-08 · ICCV 2023 1 · Yujin Jeong, Wonjeong Ryoo, SeungHyun Lee, Dabin Seo 외

In recent years, video generation has become a prominent generative tool and has drawn significant attention. However, there is little consideration in audio-to-video generation, though audio contains unique qualities li…

Video Generation

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation

2026-05-09 · Shihao Cheng, Jiaxu Zhang, Quanyue Song, Shansong Liu 외 arxiv

Motion, speech, and sound effects are fundamental elements of human-centric videos, yet their heterogeneous temporal characteristics make joint generation highly challenging. Existing audio-video generation models often …

Video Generation

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance

2025-06-26 · Akio Hayakawa, Masato Ishii, Takashi Shibuya, Yuki Mitsufuji

We propose a novel step-by-step video-to-audio generation method that sequentially produces individual audio tracks, each corresponding to a specific sound event in the video. Our approach mirrors traditional Foley workf…

Audio GenerationAudio SynthesisNegation

Video-Guided Foley Sound Generation with Multimodal Controls

2024-11-26 · CVPR 2025 1 · Ziyang Chen, Prem Seetharaman, Bryan Russell, Oriol Nieto 외

Generating sound effects for videos often requires creating artistic sound effects that diverge significantly from real-life sources and flexible control in the sound design. To address this problem, we introduce MultiFo…

Audio Generation