paper-with-me

홈 › Papers

ActionParty: Multi-Subject Action Binding in Generative Video Games

2026-04-02 · Alexander Pondaven, Ziyi Wu, Igor Gilitschenski, Philip Torr, Sergey Tulyakov, Fabio Pizzati, Aliaksandr Siarohin arxiv

Recent advances in video diffusion have enabled the development of "world models" capable of simulating interactive environments. However, these models are largely restricted to single-agent settings, failing to control multiple agents simultaneously in a scene. In this work, we tackle a fundamental issue of action binding in existing video diffusion models, which struggle to associate specific actions with their corresponding subjects. For this purpose, we propose ActionParty, an action controllable multi-subject world model for generative video games. It introduces subject state tokens, i.e. latent variables that persistently capture the state of each subject in the scene. By jointly modeling state tokens and video latents with a spatial biasing mechanism, we disentangle global video frame rendering from individual action-controlled subject updates. We evaluate ActionParty on the Melting Pot benchmark, demonstrating the first video world model capable of controlling up to seven players simultaneously across 46 diverse environments. Our results show significant improvements in action-following accuracy and identity consistency, while enabling robust autoregressive tracking of subjects through complex interactions.

📄 PDF Abstract BibTeX arXiv:2604.02330

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DisenStudio: Customized Multi-subject Text-to-Video Generation with Disentangled Spatial Control

2024-05-21 · Hong Chen, Xin Wang, YiPeng Zhang, Yuwei Zhou 외

Generating customized content in videos has received increasing attention recently. However, existing works primarily focus on customized text-to-video generation for single subject, suffering from subject-missing and at…

AttributeMotion GenerationText-to-Video GenerationVideo Generation

CogCanvas: A Benchmark for Evaluating Multi-Subject Reference-Based Image Generation

2026-06-14 · Long-Bao Nguyen, Quang-Khai Tran, Tam V. Nguyen, Minh-Triet Tran 외 arxiv

Multi-subject reference-based image generation requires jointly preserving multiple human identities, binding per-person objects and fashion items, and respecting a specified background scene, a regime where current diff…

Image Generation

T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation

2024-07-19 · CVPR 2025 1 · Kaiyue Sun, Kaiyi Huang, Xian Liu, Yue Wu 외

Text-to-video (T2V) generative models have advanced significantly, yet their ability to compose different objects, attributes, actions, and motions into a video remains unexplored. Previous text-to-video benchmarks also …

AttributeLanguage ModelingLanguage ModellingLarge Language Model+3

Generalized Protein Pocket Generation with Prior-Informed Flow Matching

2024-09-29 · Zaixi Zhang, Marinka Zitnik, Qi Liu

Designing ligand-binding proteins, such as enzymes and biosensors, is essential in bioengineering and protein biology. One critical step in this process involves designing protein pockets, the protein interface binding w…

valid

MultiBind: A Benchmark for Attribute Misbinding in Multi-Subject Generation

2026-03-23 · Wenqing Tian, Hanyi Mao, Zhaocheng Liu, Lihua Zhang 외 arxiv

Subject-driven image generation is increasingly expected to support fine-grained control over multiple entities within a single image. In multi-reference workflows, users may provide several subject images, a background …

Image Generation