paper-with-me

홈 › Papers

ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models

2026-08-14 · Xinye Li, Lingshuai Lin, Lei Wang, Liuzhou Zhang, Jialin Cui, Qingshan Li, Guanchu Wang, Qingbin Liu, Xi Chen, Jiang Bian, Wai Lam arxiv

Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step video synthesis, extending it to interactive world models remains challenging, as discrete keyboard states and continuous mouse motion must remain aligned with temporally compressed latent chunks during causal training and autoregressive rollout. We introduce ForgeWM, a progressive framework that transforms a bidirectional action-conditioned video generator into efficient few-step world models through domain adaptation, teacher-forced causal training, causal consistency distillation, and on-policy distribution matching with a bidirectional teacher. The resulting budget-specialized students operate at steady-state denoising budgets of 1, 2, and 4 steps. ForgeWM further supports a dual-path deployment protocol combining latency-critical interaction with optional replay-time refinement, where the one-step student re-noises and refines its saved draft. On paired Minecraft trajectories, ForgeWM leads the evaluated systems in Imaging Quality, reference-aligned motion-profile agreement, action-sign accuracy, and mouse-control accuracy, while achieving the lowest reference LPIPS; the same four-stage recipe transfers to gamepad-controlled FPS gameplay. Replay-time refinement matches four-step reference quality while remaining roughly three times closer to the experienced trajectory than regeneration from noise. These results demonstrate ForgeWM's effectiveness for controllable few-step video generation.

📄 PDF Abstract BibTeX arXiv:2608.14022

Code (0)

등록된 구현이 없습니다.

Tasks

Domain AdaptationVideo Generation

Similar Papers 제목 키워드 기반

PLOT: Progressive Localization via Optimal Transport in Neural Causal Abstraction

2026-05-07 · Jonathn Chang, Arya Datla, Ziv Goldfeld arxiv

Causal abstraction offers a principled framework for mechanistic interpretability, aligning a high-level causal model with the low-level computation realized by a neural network through counterfactual intervention analys…

CookGAN: Causality Based Text-to-Image Synthesis

2020-06-01 · CVPR 2020 6 · Bin Zhu, Chong-Wah Ngo

This paper addresses the problem of text-to-image synthesis from a new perspective, i.e., the cause-and-effect chain in image generation. Causality is a common phenomenon in cooking. The dish appearance changes depending…

Image Generation

Out-of-Distribution Generalization on Graphs via Progressive Inference

2025-03-04 · Yiming Xu, Bin Shi, Zhen Peng, Huixiang Liu 외

The development and evaluation of graph neural networks (GNNs) generally follow the independent and identically distributed (i.i.d.) assumption. Yet this assumption is often untenable in practice due to the uncontrollabl…

counterfactualOut-of-Distribution Generalization

SketchMamba: A Lightweight State-Space Model for Joint Progressive Sketch Classification and Stroke Auto-Completion

2026-07-26 · Kavish Jhaveri, Arya Shah arxiv

Existing vector-sketch models treat recognition and generation as separate tasks, leaving a gap for streaming interfaces that must understand a drawing as it is being made. We present SketchMamba, a single causal sequenc…

Riemann-1.0: An Embodied World Action Model for Physical AI

2026-08-27 · Haofeng Sun, Jiangbo Pei, Fei Kang, Zexiang Liu 외 arxiv

We introduce Riemann-1.0, a fully causal autoregressive World Action Model for embodied intelligence. Riemann-1.0 jointly models multi-view visual observations, robot states, and embodiment-specific actions within a unif…

Robot Manipulation