paper-with-me

Papers

Extending Visual Dynamics for Video-to-Music Generation

2025-04-10 · Xiaohao Liu, Teng Tu, Yunshan Ma, Tat-Seng Chua

Music profoundly enhances video production by improving quality, engagement, and emotional resonance, sparking growing interest in video-to-music generation. Despite recent advances, existing approaches remain limited in specific scenarios or undervalue the visual dynamics. To address these limitations, we focus on tackling the complexity of dynamics and resolving temporal misalignment between video and music representations. To this end, we propose DyViM, a novel framework to enhance dynamics modeling for video-to-music generation. Specifically, we extract frame-wise dynamics features via a simplified motion encoder inherited from optical flow methods, followed by a self-attention module for aggregation within frames. These dynamic features are then incorporated to extend existing music tokens for temporal alignment. Additionally, high-level semantics are conveyed through a cross-attention mechanism, and an annealing tuning strategy benefits to fine-tune well-trained music decoders efficiently, therefore facilitating seamless adaptation. Extensive experiments demonstrate DyViM's superiority over state-of-the-art (SOTA) methods.

📄 PDF Abstract BibTeX arXiv:2504.07594

Code (0)

등록된 구현이 없습니다.

Tasks

Music GenerationOptical Flow Estimation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization

2024-10-16 · RuiQi Li, Siqi Zheng, Xize Cheng, Ziang Zhang 외

Generating music that aligns with the visual content of a video has been a challenging task, as it requires a deep understanding of visual semantics and involves generating music whose melody, rhythm, and dynamics harmon…

In-Context LearningMusic GenerationRhythm

KARMA-MV: A Benchmark for Causal Question Answering on Music Videos

2026-05-05 · Archishman Ghosh, Abhinaba Roy, Dorien Herremans arxiv

While significant progress has been made in Video Question Answering and cross-modal understanding, causal reasoning about how visual dynamics drive musical structure in music videos remains under-explored. We introduce …

Video Question Answering

V2Meow: Meowing to the Visual Beat via Video-to-Music Generation

2023-05-11 · Kun Su, Judith Yue Li, Qingqing Huang, Dima Kuzmin 외

Video-to-music generation demands both a temporally localized high-quality listening experience and globally aligned video-acoustic signatures. While recent music generation models excel at the former through advanced au…

Music Generation

Customized Condition Controllable Generation for Video Soundtrack

2025-01-01 · CVPR 2025 1 · Fan Qi, Kunsheng Ma, Changsheng Xu

Recent advances in latent diffusion models (LDMs) have enabled data-driven paradigms for video soundtrack generation, improving multimodal alignment capabilities. However, current two-stage frameworks--which separate…

Audio Synthesis

OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data

2026-06-29 · Kaixing Yang, Jiashu Zhu, Xulong Tang, Ziqiao Peng 외 arxiv

Music-driven dance video generation aims to synthesize expressive human motion that is temporally aligned with music while maintaining high visual fidelity. Despite recent progress, existing methods still face two key li…

Video Generation