paper-with-me

Papers

Mem-World: Memory-Augmented Action-Conditioned World Models for Persistent Robot Manipulation

2026-06-17 · Zirui Zheng, Jiaqian Yu, Xiongfeng Peng, jun shi, Mingyi Li, Chao Zhang, Weiming Li, Dong Wang, Huchuan Lu, Xu Jia arxiv

Action-conditioned world models have emerged as a promising paradigm for robot learning, offering a scalable alternative to costly real-world experimentation by generating action-consistent video rollouts. However, persistent world modeling remains challenging in manipulation: frequent end-effector occlusions and rapid wrist-camera motion make the current observation insufficient for predicting future views, causing models to forget or hallucinate scene details seen in earlier frames. Existing memory retrieval strategies often fail to identify informative history in dynamic manipulation scenarios. To address this limitation, we propose Mem-World, a memory-augmented multi-view action-conditioned world model. At its core, we present W-VMem, a 4D wrist-view-centered surfel-indexed memory that anchors historical observations to temporally evolving surface elements. By explicitly modeling when and where scene elements are observed, W-VMem enables geometry-aware retrieval of relevant history frames conditioned on future actions. During generation, relevant history frames are selected via surfel-based rendering and scoring, providing informative and non-redundant context for prediction. Extensive experiments show that Mem-World generates persistent rollouts in complex manipulation scenarios, enables more reliable policy evaluation than Ctrl-World, improving the Pearson correlation with real-world performance by 14.5\%, and supports effective policy improvement through synthetic data generation, increasing success rates from 58\% to 72\% on long-horizon tasks.

📄 PDF Abstract BibTeX arXiv:2606.18960

Code (0)

등록된 구현이 없습니다.

Tasks

Synthetic Data GenerationRobot Manipulation

Similar Papers 제목 키워드 기반

MemWM: Memory-Augmented Text-Based World Model

2026-08-07 · Yujun Wang, Tao Zhang, Jinhe Bi, Aniri 외 arxiv

World models are increasingly used to support planning in agents by predicting how environment states evolve in response to agent actions. Yet fluent next-state predictions can still omit task-critical facts, corrupt pro…

BamaER: A Behavior-Aware Memory-Augmented Model for Exercise Recommendation

2026-02-03 · Qing Yang, Yuhao Jiang, Rui Wang, Jipeng Guo 외 arxiv

Exercise recommendation focuses on personalized exercise selection conditioned on students' learning history, personal interests, and other individualized characteristics. Despite notable progress, most existing methods …

Knowledge Tracing

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

2026-07-21 · Fan Jiang, Zhaoxu Sun, Mengchao Wang, Ziyu Zhu 외 hf

We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet v…

WorldKV: Efficient World Memory with World Retrieval and Compression

2026-05-21 · Jung Yi, Minjae Kim, Paul Hyunbin Cho, Wooseok Jang 외 arxiv

Autoregressive video diffusion models have enabled real-time, action-conditioned world generation. However, sustaining a persistent world, where revisiting a previously seen viewpoint yields consistent content, remains a…

MemoryWAM: Efficient World Action Modeling with Persistent Memory

2026-06-18 · Sizhe Yang, Juncheng Mu, Tianming Wei, Chenhao Lu 외 arxiv

Robust robotic manipulation in the real world requires not only an understanding of the current observation, but also memory and dynamics modeling. World action models (WAMs) possess these capabilities by jointly modelin…

Computational Efficiency