paper-with-me

Papers

Plan-X: Instruct Video Generation via Semantic Planning

2025-11-22 · Lun Huang, You Xie, Hongyi Xu, Tianpei Gu, Chenxu Zhang, Guoxian Song, Zenan Li, Xiaochen Zhao, Linjie Luo, Guillermo Sapiro arxiv

Diffusion Transformers have demonstrated remarkable capabilities in visual synthesis, yet they often struggle with high-level semantic reasoning and long-horizon planning. This limitation frequently leads to visual hallucinations and mis-alignments with user instructions, especially in scenarios involving complex scene understanding, human-object interactions, multi-stage actions, and in-context motion reasoning. To address these challenges, we propose Plan-X, a framework that explicitly enforces high-level semantic planning to instruct video generation process. At its core lies a Semantic Planner, a learnable multimodal language model that reasons over the user's intent from both text prompts and visual context, and autoregressively generates a sequence of text-grounded spatio-temporal semantic tokens. These semantic tokens, complementary to high-level text prompt guidance, serve as structured "semantic sketches" over time for the video diffusion model, which has its strength at synthesizing high-fidelity visual details. Plan-X effectively integrates the strength of language models in multimodal in-context reasoning and planning, together with the strength of diffusion models in photorealistic video synthesis. Extensive experiments demonstrate that our framework substantially reduces visual hallucinations and enables fine-grained, instruction-aligned video generation consistent with multimodal context.

📄 PDF Abstract BibTeX arXiv:2511.17986

Code (0)

등록된 구현이 없습니다.

Tasks

Scene UnderstandingVideo Generation

Similar Papers 제목 키워드 기반

This&That: Language-Gesture Controlled Video Generation for Robot Planning

2024-07-08 · Boyang Wang, Nikhil Sridhar, Chao Feng, Mark Van der Merwe 외

Clear, interpretable instructions are invaluable when attempting any complex task. Good instructions help to clarify the task and even anticipate the steps needed to solve it. In this work, we propose a robot learning fr…

Task PlanningVideo Generation

Planning with Sketch-Guided Verification for Physics-Aware Video Generation

2025-11-21 · Yidong Huang, Zun Wang, Han Lin, Dong-Ki Kim 외 arxiv

Recent video generation approaches increasingly rely on planning intermediate control signals such as object trajectories to improve temporal coherence and motion fidelity. However, these methods mostly employ single-sho…

Video GenerationMotion Planning

Event-Guided Procedure Planning from Instructional Videos with Text Supervision

2023-08-17 · ICCV 2023 1 · An-Lan Wang, Kun-Yu Lin, Jia-Run Du, Jingke Meng 외

In this work, we focus on the task of procedure planning from instructional videos with text supervision, where a model aims to predict an action sequence to transform the initial visual state into the goal visual state.…

UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving

2025-12-10 · Hao Lu, Ziyang Liu, Guangfeng Jiang, Yuanfei Luo 외 arxiv

Autonomous driving (AD) systems struggle in long-tail scenarios due to limited world knowledge and weak visual dynamic modeling. Existing vision-language-action (VLA)-based methods cannot leverage unlabeled videos for vi…

Trajectory PlanningAutonomous DrivingVideo Generation

Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation

2025-12-02 · Zeqi Xiao, Yiwei Zhao, Lingxiao Li, Yushi Lan 외 arxiv

We investigate whether video generative models can exhibit visuospatial intelligence, a capability central to human cognition, using only visual data. To this end, we present Video4Spatial, a framework showing that video…

Instruction FollowingVideo Generation