paper-with-me

Papers

Vidarc: Embodied Video Diffusion Model for Closed-loop Control

2025-12-19 · Yao Feng, Chendong Xiang, Xinyi Mao, Hengkai Tan, Zuyue Zhang, Shuhe Huang, Kaiwen Zheng, Haitian Liu, Hang Su, Jun Zhu arxiv

Robotic arm manipulation in data-scarce settings is a highly challenging task due to the complex embodiment dynamics and diverse contexts. Recent video-based approaches have shown great promise in capturing and transferring the temporal and physical interactions by pre-training on Internet-scale video data. However, such methods are often not optimized for the embodiment-specific closed-loop control, typically suffering from high latency and insufficient grounding. In this paper, we present Vidarc (Video Diffusion for Action Reasoning and Closed-loop Control), a novel autoregressive embodied video diffusion approach augmented by a masked inverse dynamics model. By grounding video predictions with action-relevant masks and incorporating real-time feedback through cached autoregressive generation, Vidarc achieves fast, accurate closed-loop control. Pre-trained on one million cross-embodiment episodes, Vidarc surpasses state-of-the-art baselines, achieving at least a 15% higher success rate in real-world deployment and a 91% reduction in latency. We also highlight its robust generalization and error correction capabilities across previously unseen robotic platforms.

📄 PDF Abstract BibTeX arXiv:2512.17661

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

World-in-World: World Models in a Closed-Loop World

2025-10-20 · Jiahan Zhang, Muqing Jiang, Nanru Dai, Taiming Lu 외 arxiv

Generative world models (WMs) can now simulate worlds with striking visual realism, which naturally raises the question of whether they can endow embodied agents with predictive perception for decision making. Progress o…

Decision Making

KineBench: Benchmarking Embodied World Models via IDM-Free Kinematic Grounding

2026-07-22 · Zeyu Liu, Zhangzhe Zhu, Yang Zhang, Chenyou Fan 외 arxiv

Evaluating the physical consistency of embodied world models(EWMs) is a critical open challenge. While closed-loop evaluation via simulator rollouts offers a more faithful assessment of physical plausibility than open-lo…

Video Generation

SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios

2025-11-20 · Juntao Cheng, Wanyue Zhang, Zhiwei Yu, Shuo Ren 외 arxiv

Tangible control interfaces (TCIs), such as appliance panels, remotes, elevators, and embedded GUIs, are a fundamental component of everyday human-built environments. Interacting with these interfaces requires agents not…

Video Generation

StreamingClaw Technical Report

2026-03-23 · Jiawei Chen, Zhe Chen, Chaoqun Du, Maokui He 외 arxiv

Emerging applications such as embodied intelligence, AI hardware, autonomous driving, and intelligent cockpits rely on a real-time perception-decision-action closed loop, posing stringent challenges for streaming video u…

Autonomous Driving

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind

2026-06-04 · Jianxu Shangguan, Jing Xu, Hang Ye, Xiaoxuan Ma 외 arxiv

Creating lifelike digital humans with genuine social intelligence requires unifying cognitive reasoning and multimodal generation within a coherent framework. Current approaches treat these as separate tasks: Large Langu…

multimodal generationVideo Generation