paper-with-me

Papers

Controllable Video-to-Music Generation with Multiple Time-Varying Conditions

2025-07-28 · Junxian Wu, Weitao You, Heda Zuo, Dengming Zhang, Pei Chen, Lingyun Sun arxiv

Music enhances video narratives and emotions, driving demand for automatic video-to-music (V2M) generation. However, existing V2M methods relying solely on visual features or supplementary textual inputs generate music in a black-box manner, often failing to meet user expectations. To address this challenge, we propose a novel multi-condition guided V2M generation framework that incorporates multiple time-varying conditions for enhanced control over music generation. Our method uses a two-stage training strategy that enables learning of V2M fundamentals and audiovisual temporal synchronization while meeting users' needs for multi-condition control. In the first stage, we introduce a fine-grained feature selection module and a progressive temporal alignment attention mechanism to ensure flexible feature alignment. For the second stage, we develop a dynamic conditional fusion module and a control-guided decoder module to integrate multiple conditions and accurately guide the music composition process. Extensive experiments demonstrate that our method outperforms existing V2M pipelines in both subjective and objective evaluations, significantly enhancing control and alignment with user expectations.

📄 PDF Abstract BibTeX arXiv:2507.20627

Code (0)

등록된 구현이 없습니다.

Tasks

Music Generation

Similar Papers 제목 키워드 기반

V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation

2026-03-11 · Yan-Bo Lin, Jonah Casebeer, Long Mai, Aniruddha Mahapatra 외 arxiv

Generating music that temporally aligns with video events is challenging for existing text-to-music models, which lack fine-grained temporal control. We introduce V2M-ZERO, a video-to-music generation approach that gener…

Music Generation

ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion

2025-07-26 · Xuanchen Wang, Heng Wang, Weidong Cai arxiv

Modern artistic productions increasingly demand automated choreography generation that adapts to diverse musical styles and individual dancer characteristics. Existing approaches often fail to produce high-quality dance …

Video GenerationStyle Transfer

Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation

2024-12-12 · Baisen Wang, Le Zhuo, Zhaokai Wang, Chenxi Bao 외

Multimodal music generation aims to produce music from diverse input modalities, including text, videos, and images. Existing methods use a common embedding space for multimodal fusion. Despite their effectiveness in oth…

cross-modal alignmentMultimodal Music GenerationMusic GenerationRetrieval

Video Object Segmentation-Aware Audio Generation

2025-09-30 · Ilpo Viertola, Vladimir Iashin, Esa Rahtu arxiv

Existing multimodal audio generation models often lack precise user control, which limits their applicability in professional Foley workflows. In particular, these models focus on the entire video and do not provide prec…

Video Object SegmentationAudio Generation

XMusic: Towards a Generalized and Controllable Symbolic Music Generation Framework

2025-01-15 · Sida Tian, Can Zhang, Wei Yuan, Wei Tan 외

In recent years, remarkable advancements in artificial intelligence-generated content (AIGC) have been achieved in the fields of image synthesis and text generation, generating content comparable to that produced by huma…

Emotion RecognitionImage GenerationMulti-Task LearningMusic Generation+1