paper-with-me

Papers

GraphVid: Interactive Graph-Controllable Video Generation

2026-07-23 · Vedant Shah, Onkar Susladkar, Tushar Prakash, Kiet Nguyen, Tianjio Yu, Adheesh Juvekar, Muntasir Waheed, Ismini Lourentzou arxiv

Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to draw accurate tracks for multiple objects, which scales poorly with scene complexity and becomes ambiguous under occlusion or overlap. To enable flexible yet precise multi-subject control, we introduce GraphVid, a graph-conditioned image-to-video generation model that enables interactive control through structured interaction graphs. We further curate GraphVid-Bench, a large-scale interaction-centric video dataset with structured relational annotations to enable training of interaction-aware video generation models. Despite using substantially less training data and fewer trainable parameters than prior motion-control methods, GraphVid delivers strong controllability and video quality. Compared with Motion-I2V, GraphVid reduces FID by up to 39.9% and FVD by 37.6%, while improving PSNR (9.87=>15.98) and SSIM (0.38=>0.61). Our results highlight the potential of structured semantic interfaces as a powerful paradigm for controllable video generation.

📄 PDF Abstract BibTeX arXiv:2607.21580

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Understanding Long Videos via LLM-Powered Entity Relation Graphs

2025-01-27 · Meng Chu, Yicong Li, Tat-Seng Chua

The analysis of extended video content poses unique challenges in artificial intelligence, particularly when dealing with the complexity of tracking and understanding visual elements across time. Current methodologies th…

EgoSchemaLarge Language ModelObject TrackingRelation+1

Click2Graph: Interactive Panoptic Video Scene Graphs from a Single Click

2025-11-20 · Raphael Ruschel, Hardikkumar Prajapati, Awsafur Rahman, B. S. Manjunath arxiv

State-of-the-art Video Scene Graph Generation (VSGG) systems provide structured visual understanding but operate as closed, feed-forward pipelines with no ability to incorporate human guidance. In contrast, promptable se…

Video scene graph generationRelational ReasoningScene Understanding

Matrix-Game: Interactive World Foundation Model

2025-06-23 · Yifan Zhang, Chunli Peng, Boyang Wang, Puyi Wang 외

We introduce Matrix-Game, an interactive world foundation model for controllable game world generation. Matrix-Game is trained using a two-stage pipeline that first performs large-scale unlabeled pretraining for environm…

MinecraftmodelVideo Generation

InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions

2024-02-05 · Yiyuan Zhang, Yuhao Kang, Zhixin Zhang, Xiaohan Ding 외

We introduce $\textit{InteractiveVideo}$, a user-centric framework for video generation. Different from traditional generative approaches that operate based on user-provided images or text, our framework is designed for …

Video Generation

Yan: Foundational Interactive Video Generation

2025-08-12 · Deheng Ye, Fangyun Zhou, Jiacheng Lv, Jianqi Ma 외 arxiv

We present Yan, a foundational framework for interactive video generation, covering the entire pipeline from simulation and generation to editing. Specifically, Yan comprises three core modules. AAA-level Simulation: We …

Video Generation