paper-with-me

Papers

Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation

2026-04-19 · Vaibhavi Lokegaonkar, Aryan Vijay Bhosale, Vishnu Raj, Gouthaman KV, Ramani Duraiswami, Lie Lu, Sreyan Ghosh, Dinesh Manocha arxiv

Video-to-music (V2M) is the fundamental task of creating background music for an input video. Recent V2M models achieve audiovisual alignment by typically relying on visual conditioning alone and provide limited semantic and stylistic controllability to the end user. In this paper, we present Video-Robin, a novel text-conditioned video-to-music generation model that enables fast, high-quality, semantically aligned music generation for video content. To balance musical fidelity and semantic understanding, Video-Robin integrates autoregressive planning with diffusion-based synthesis. Specifically, an autoregressive module models global structure by semantically aligning visual and textual inputs to produce high-level music latents. These latents are subsequently refined into coherent, high-fidelity music using local Diffusion Transformers. By factoring semantically driven planning into diffusion-based synthesis, Video-Robin enables fine-grained creator control without sacrificing audio realism. Our proposed model outperforms baselines that solely accept video input and additional feature conditioned baselines on both in-distribution and out-of-distribution benchmarks with a 2.21x speed in inference compared to SOTA. We will open-source everything upon paper acceptance.

📄 PDF Abstract BibTeX arXiv:2604.17656

Code (0)

등록된 구현이 없습니다.

Tasks

Music Generation

Similar Papers 제목 키워드 기반

Plan-X: Instruct Video Generation via Semantic Planning

2025-11-22 · Lun Huang, You Xie, Hongyi Xu, Tianpei Gu 외 arxiv

Diffusion Transformers have demonstrated remarkable capabilities in visual synthesis, yet they often struggle with high-level semantic reasoning and long-horizon planning. This limitation frequently leads to visual hallu…

Scene UnderstandingVideo Generation

Macro-from-Micro Planning for High-Quality and Parallelized Autoregressive Long Video Generation

2025-08-05 · Xunzhi Xiang, Yabo Chen, Guiyu Zhang, Zhongyu Wang 외 arxiv

Current autoregressive diffusion models excel at video generation but are generally limited to short temporal durations. Our theoretical analysis indicates that the autoregressive modeling typically suffers from temporal…

Video Generation

Epona: Autoregressive Diffusion World Model for Autonomous Driving

2025-06-30 · Kaiwen Zhang, Zhenyu Tang, Xiaotao Hu, Xingang Pan 외

Diffusion models have demonstrated exceptional visual quality in video generation, making them promising for autonomous driving world modeling. However, existing video diffusion-based world models struggle with flexible-…

Autonomous DrivingmodelMotion PlanningNavSim+3

AID: Agent Intent from Diffusion for Multi-Agent Informative Path Planning

2025-12-02 · Jeric Lew, Yuhong Cao, Derek Ming Siang Tan, Guillaume Sartoretti arxiv

Information gathering in large-scale or time-critical scenarios (e.g., environmental monitoring, search and rescue) requires broad coverage within limited time budgets, motivating the use of multi-agent systems. These sc…

Reinforcement Learning

MarDini: Masked Autoregressive Diffusion for Video Generation at Scale

2024-10-26 · Haozhe Liu, Shikun Liu, Zijian Zhou, Mengmeng Xu 외

We introduce MarDini, a new family of video diffusion models that integrate the advantages of masked auto-regression (MAR) into a unified diffusion model (DM) framework. Here, MAR handles temporal planning, while DM focu…

Image to Video GenerationVideo Generation