paper-with-me

Papers

Customized Condition Controllable Generation for Video Soundtrack

2025-01-01 · CVPR 2025 1 · Fan Qi, Kunsheng Ma, Changsheng Xu

Recent advances in latent diffusion models (LDMs) have enabled data-driven paradigms for video soundtrack generation, improving multimodal alignment capabilities. However, current two-stage frameworks--which separately optimize audio-visual correspondence and conditional audio synthesis--fundamentally limit joint modeling of dynamic acoustic properties. In this paper, we propose a novel framework for generating video soundtracks that simultaneously produces music and sound effect tailored to the video content. Our method incorporates a Contrastive Visual-Sound-Music pretraining process that maps these modalities into a unified feature space, enhancing the model's ability to capture intricate audio dynamics. We design Spectrum Divergence Masked Attention for Unet to differentiate between the unique characteristics of sound effect and music. We utilize Score-guided Noise Iterative Optimization to provide musicians with customizable control during the generation process. Extensive evaluations on the FilmScoreDB and SymMV&HIMV datasets demonstrate that our approach significantly outperforms state-of-the-art baselines in both subjective and objective assessments, highlighting its potential as a robust tool for video soundtrack generation.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Synthesis

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Long-Term Rhythmic Video Soundtracker

2023-05-02 · Jiashuo Yu, Yaohui Wang, Xinyuan Chen, Xiao Sun 외

We consider the problem of generating musical soundtracks in sync with rhythmic visual cues. Most existing works rely on pre-defined music representations, leading to the incompetence of generative flexibility and comple…

HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

2025-05-07 · Teng Hu, Zhentao Yu, Zhengguang Zhou, Sen Liang 외

Customized video generation aims to produce videos featuring specific subjects under flexible user-defined conditions, yet existing methods often struggle with identity consistency and limited input modalities. In this p…

Human-Domain Subject-to-VideoSingle-Domain Subject-to-VideoVideo AlignmentVideo Generation

Conditional Generation of Audio from Video via Foley Analogies

2023-04-17 · CVPR 2023 1 · Yuexi Du, Ziyang Chen, Justin Salamon, Bryan Russell 외

The sound effects that designers add to videos are designed to convey a particular artistic effect and, thus, may be quite different from a scene's true sound. Inspired by the challenges of creating a soundtrack for a vi…

Automated Composition of Picture-Synched Music Soundtracks for Movies

2019-10-19 · Vansh Dassani, Jon Bird, Dave Cliff

We describe the implementation of and early results from a system that automatically composes picture-synched musical soundtracks for videos and movies. We use the phrase "picture-synched" to mean that the structure of t…

Music Generation

The NES Video-Music Database: A Dataset of Symbolic Video Game Music Paired with Gameplay Videos

2024-04-05 · Igor Cardoso, Rubens O. Moraes, Lucas N. Ferreira

Neural models are one of the most popular approaches for music generation, yet there aren't standard large datasets tailored for learning music directly from game data. To address this research gap, we introduce a novel …

Music Generation