paper-with-me

홈 › Papers

EWAM: An Enhanced World Action Model for Closed-Loop Online Adaptation in Embodied Intelligence

2026-06-10 · Xin Zhou, Cong Miao arxiv

In this paper, we propose the Enhanced World Action Model (EWAM), a closed-loop online adaptation architecture built upon a pretrained and fully frozen Cosmos3 backbone network. Evaluated entirely under a zero-shot task protocol, EWAM is centrally focused on reducing the amount of additional deployment data required to adapt to new task layouts. Notably, no extra task-specific demonstration sets were introduced in any of the evaluations, and no fine-tuning was performed on the backbone network. Its performance gains stem entirely from an inference-time co-reasoning mechanism composed of four inserted lightweight neural layers: the Neural Experience Memory Layer located in the intermediate layers of the Diffusion Transformer (DiT) provides task-relevant execution context; the Neural Anomaly Detection Layer after the state prediction head monitors the divergence between predicted and actual states in real time; the Neural Policy Routing Layer dynamically selects direct execution, conservative replanning, or rollback recovery based on the anomaly severity; and the Neural Action Correction Layer refines the generated action chunks using execution diagnostics. Unlike naive feature fusion, the memory, anomaly detection, and correction modules are deeply integrated into the Cosmos3 forward path in a differentiable manner, with only the final routing decision being a discrete supervised one.

📄 PDF Abstract BibTeX arXiv:2606.12690

Code (0)

등록된 구현이 없습니다.

Tasks

Anomaly Detection

Similar Papers 제목 키워드 기반

GameWAM: A World Action Model for Video Games

2026-08-25 · Yuncheng Guo, Zhanqiu Zhang, Yiwen Guo, Weijia Li hf

Modern video games combine first-person perception, rapid visual changes, persistent world state, and heterogeneous native controls. Existing game agents map visual and task context directly to actions but lack explicit …

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

2026-06-17 · Yuyang Zhang, Wenyao Zhang, Zekun Qi, He Zhang 외 arxiv

World Action Models (WAMs) commonly rely on video generation to bridge visual world modeling and robot control. However, video-based WAMs face three coupled limitations: dense multi-frame future tokens make inference cos…

Video PredictionVideo GenerationImage Editing

DriveWAM: Video Generative Priors Enable Scalable World-Action Modeling for Autonomous Driving

2026-05-27 · Chen Shi, Jinrui Xu, Shaoshuai Shi, Kehua Sheng 외 arxiv

Pretrained foundation models have become an important basis for end-to-end autonomous driving. In contrast to vision-language models pretrained primarily on static image-text pairs, video generative models capture tempor…

Scene UnderstandingAutonomous Driving

StageWAM: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation

2026-08-11 · Xiao Liu, Yuguang Yang, Xi Wang, Kai Jiang 외 arxiv

Generalist robot policies aim to map multimodal observations and linguistic task instructions to actions across diverse tasks. However, existing methods typically represent the future as a fixed, short video-action chunk…

Robot Manipulation

MobileWAM: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight

2026-08-05 · Zehua Fan, Junjie He, Wenxuan Song, Xi Wang 외 arxiv

World action models (WAMs) built on video generation backbones are a rising recipe for robot learning, yet remain confined to tabletop manipulation. Mobile manipulation demands simultaneous locomotion and whole-body mani…

Video Generation