paper-with-me

홈 › Papers

Soundini: Sound-Guided Diffusion for Natural Video Editing

2023-04-13 · Seung Hyun Lee, Sieun Kim, Innfarn Yoo, Feng Yang, Donghyeon Cho, Youngseo Kim, Huiwen Chang, Jinkyu Kim, Sangpil Kim

We propose a method for adding sound-guided visual effects to specific regions of videos with a zero-shot setting. Animating the appearance of the visual effect is challenging because each frame of the edited video should have visual changes while maintaining temporal consistency. Moreover, existing video editing solutions focus on temporal consistency across frames, ignoring the visual style variations over time, e.g., thunderstorm, wave, fire crackling. To overcome this limitation, we utilize temporal sound features for the dynamic style. Specifically, we guide denoising diffusion probabilistic models with an audio latent representation in the audio-visual latent space. To the best of our knowledge, our work is the first to explore sound-guided natural video editing from various sound sources with sound-specialized properties, such as intensity, timbre, and volume. Additionally, we design optical flow-based guidance to generate temporally consistent video frames, capturing the pixel-wise relationship between adjacent frames. Experimental results show that our method outperforms existing video editing techniques, producing more realistic visual effects that reflect the properties of sound. Please visit our page: https://kuai-lab.github.io/soundini-gallery/.

📄 PDF Abstract BibTeX arXiv:2304.06818

Code (1)

kuai-lab/soundini-official 공식 구현

Tasks

DenoisingOptical Flow EstimationVideo Editing

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

The Power of Sound (TPoS): Audio Reactive Video Generation with Stable Diffusion

2023-09-08 · ICCV 2023 1 · Yujin Jeong, Wonjeong Ryoo, SeungHyun Lee, Dabin Seo 외

In recent years, video generation has become a prominent generative tool and has drawn significant attention. However, there is little consideration in audio-to-video generation, though audio contains unique qualities li…

Video Generation

Label-free Motion-Conditioned Diffusion Model for Cardiac Ultrasound Synthesis

2025-12-10 · Zhe Li, Hadrien Reynaud, Johanna P Müller, Bernhard Kainz arxiv

Ultrasound echocardiography is essential for the non-invasive, real-time assessment of cardiac function, but the scarcity of labelled data, driven by privacy restrictions and the complexity of expert annotation, remains …

Video Generation

Images that Sound: Composing Images and Sounds on a Single Canvas

2024-05-20 · Ziyang Chen, Daniel Geng, Andrew Owens

Spectrograms are 2D representations of sound that look very different from the images found in our visual world. And natural images, when played as spectrograms, make unnatural sounds. In this paper, we show that it is p…

Align, Adapt and Inject: Sound-guided Unified Image Generation

2023-06-20 · Yue Yang, Kaipeng Zhang, Yuying Ge, Wenqi Shao 외

Text-guided image generation has witnessed unprecedented progress due to the development of diffusion models. Beyond text and image, sound is a vital element within the sphere of human perception, offering vivid represen…

Image GenerationRetrievalText Retrieval

Animate and Sound an Image

2025-01-01 · CVPR 2025 1 · Xihua Wang, Ruihua Song, Chongxuan Li, Xin Cheng 외

This paper addresses a promising yet underexplored task, Image-to-Sounding-Video (I2SV) generation, which animates a static image and generates synchronized sound simultaneously. Despite advances in video and audio g…

Audio Generation