paper-with-me

홈 › Papers

RoboMME-Interference: Benchmarking Robot Memory Under Interference

2026-06-21 · Soumil Rathi arxiv

Robots deployed in realistic settings will accumulate experience across many sessions and tasks over their deployment. The robot's tasks may often require it to remember information from multiple sessions ago, making long-context robot memory important for real-world deployments. However, most robot-memory benchmarks today are based on single episodes or a short context. To measure how current robot memory systems perform on longer sessions with more distractions, we introduce RoboMME-Interference, a cross-session benchmark built on RoboMME (Dai et al., 2026). For each query episode, we construct a session history using the query's relevant prior demonstration followed by a controlled number of unrelated sessions, which we provide to the VLA as memory and measure accuracy. Running RoboMME's released memory-augmented $π_{0.5}$ variants unmodified through this benchmark, we find that while perceptual memory variants improve success when given the history without any distractors, they decay strongly and steadily as unrelated sessions accumulate. The subgoal variants, which read the history with a vision-language model and pass written subgoals to the policy, improve less at their best but hold more of that improvement as distractors accumulate. Adding a retrieval step to the strongest perceptual variant, which selects the section of history most visually similar to the robot's current view and passes only that section to the policy, restores its no-distractor success rate at every interference level. With this release, we emphasize the importance of long-context memory and robustness to interference and show that current systems largely fail on such capabilities. The project page, videos, code, and data are at https://robotmemorybench.com.

📄 PDF Abstract BibTeX arXiv:2606.22338

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies

2026-03-04 · Yinpei Dai, Hongze Fu, Jayjun Lee, Yuejiang Liu 외 arxiv

Memory is critical for long-horizon and history-dependent robotic manipulation. Such tasks often involve counting repeated actions or manipulating objects that become temporarily occluded. Recent vision-language-action (…

PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control

2026-08-25 · Suhwan Choi, Jaeyoon Jung, Sungkyung Kim, Yunsung Lee 외 hf

Multimodal large language models (MLLMs) can integrate long visual histories, reason under partial observability, and infer behavior from a few examples. Yet vision-language-action (VLA) models generally inherit pretrain…

WeaveLA: Event Driven Cross-Subtask Latent Memory Weaving for Repetitive Robot Manipulation

2026-06-16 · Shoujing Zhu, Zhenyang Liu, Fungmiu Wang, Jiafeng Wang 외 arxiv

Vision-Language-Action (VLA) policies have achieved remarkable single-step manipulation, yet they remain brittle precisely where each stage depends on what was just completed. The core issue is structural: short-window V…

Robot Manipulation

DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?

2026-06-10 · Jadelynn Dao, Milan Ganai, Yasmina Abukhadra, Ajay Sridhar 외 arxiv

Vision-Language Models (VLMs) are increasingly deployed as high-level planners for embodied agents, with an emerging strategy of scaling test-time compute to improve capability. However, we observe that doing so increase…

UTILMEM: Benchmarking Evidence Utilization in Long-Term Conversational Memory

2026-08-31 · Peijun Qing, Fobo Shi, Soroush Vosoughi arxiv

Long-term memory is increasingly important for conversational agents, yet existing benchmarks primarily measure memory through pointwise factual recall: whether a system can recover isolated facts or event-level details …