paper-with-me

Papers

Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement Learning

2025-02-14 · Egor Cherepanov, Nikita Kachaev, Alexey K. Kovalev, Aleksandr I. Panov

Memory is crucial for enabling agents to tackle complex tasks with temporal and spatial dependencies. While many reinforcement learning (RL) algorithms incorporate memory, the field lacks a universal benchmark to assess an agent's memory capabilities across diverse scenarios. This gap is particularly evident in tabletop robotic manipulation, where memory is essential for solving tasks with partial observability and ensuring robust performance, yet no standardized benchmarks exist. To address this, we introduce MIKASA (Memory-Intensive Skills Assessment Suite for Agents), a comprehensive benchmark for memory RL, with three key contributions: (1) we propose a comprehensive classification framework for memory-intensive RL tasks, (2) we collect MIKASA-Base -- a unified benchmark that enables systematic evaluation of memory-enhanced agents across diverse scenarios, and (3) we develop MIKASA-Robo (pip install mikasa-robo-suite) -- a novel benchmark of 32 carefully designed memory-intensive tasks that assess memory capabilities in tabletop robotic manipulation. Our work introduces a unified framework to advance memory RL research, enabling more robust systems for real-world use. MIKASA is available at https://tinyurl.com/membenchrobots.

📄 PDF Abstract BibTeX arXiv:2502.10550

Code (1)

CognitiveAISystems/MIKASA-Robo 공식 구현 pytorch

Tasks

Reinforcement Learning (RL)Skills Assessment

Similar Papers 제목 키워드 기반

Worth Remembering: Surprise-Gated Robot Episodic Memory

2026-06-02 · Nicolas Gorlo, Derek K. Wise, Alberto Speranzon, Luca Carlone arxiv

Robots solving generalist tasks need to be able to ground instructions in their past experience, since humans may refer to notable past events when giving a task (e.g., ``Take me to where the chemical spill happened yest…

Event SegmentationQuestion Answering

RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark

2026-05-11 · Huashuo Lei, Wenxuan Song, Huarui Zhang, Jieyuan Pei 외 arxiv

Memory is a critical component of robotic intelligence, as robots must rely on past observations and actions to accomplish long-horizon tasks in partially observable environments. However, existing robotic memory benchma…

Evolution Gym: A Large-Scale Benchmark for Evolving Soft Robots

2022-01-24 · NeurIPS 2021 12 · Jagdeep Singh Bhatia, Holly Jackson, Yunsheng Tian, Jie Xu 외

Both the design and control of a robot play equally important roles in its task performance. However, while optimal control is well studied in the machine learning and robotics community, less attention is placed on find…

Deep Reinforcement Learning

Meta-Memory: Retrieving and Integrating Semantic-Spatial Memories for Robot Spatial Reasoning

2025-09-25 · Yufan Mao, Hanjing Ye, Wenlong Dong, Chengjie Zhang 외 arxiv

Navigating complex environments requires robots to effectively store observations as memories and leverage them to answer human queries about spatial locations, which is a critical yet underexplored research challenge. W…

Spatial Reasoning

InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding

2025-06-18 · Minsoo Kim, Kyuhong Shim, Jungwook Choi, Simyung Chang

Modern multimodal large language models (MLLMs) can reason over hour-long video, yet their key-value (KV) cache grows linearly with time--quickly exceeding the fixed memory of phones, AR glasses, and edge robots. Prior c…

GPUStreaming video understandingTARVideo Understanding