paper-with-me

Papers

LongBench: Evaluating Robotic Manipulation Policies on Real-World Long-Horizon Tasks

2026-04-18 · Xueyao Chen, Jingkai Jia, Tong Yang, Yibo Fu, Wei Li, Wenqiang Zhang arxiv

Robotic manipulation policies often degrade over extended horizons, yet existing benchmarks provide limited insight into why such failures occur. Most prior benchmarks are either simulation-based or report aggregate success, making it difficult to disentangle the distinct sources of temporal difficulty in real-world execution. We introduce LongBench, a real-world benchmark for evaluating long-horizon manipulation. LongBench consists of over 1,000 real-world episodes, covering two complementary regimes: Context-Independent (fully observable) and Context-Dependent (ambiguity-driven). By organizing tasks into capability- and ambiguity-specific subsets, LongBench enables mechanism-aware evaluation of execution robustness, temporal consistency, and context-dependent reasoning. Evaluating six state-of-the-art policies reveals that long-horizon performance is not governed by a single factor. We observe that performance in fully observable settings is more strongly associated with execution robustness, while contextual difficulty varies across tasks and is not consistently improved by memory-based methods. We hope that LongBench serves as a useful benchmark for studying long-horizon manipulation and for developing policies with stronger robustness across both execution and contextual challenges.

📄 PDF Abstract BibTeX arXiv:2604.16788

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Robot Policy Evaluation for Sim-to-Real Transfer: A Benchmarking Perspective

2025-08-14 · Xuning Yang, Clemens Eppner, Jonathan Tremblay, Dieter Fox 외 arxiv

Current vision-based robotics simulation benchmarks have significantly advanced robotic manipulation research. However, robotics is fundamentally a real-world problem, and evaluation for real-world applications has lagge…

Scalable Policy Evaluation with Video World Models

2025-11-14 · Wei-Cheng Tseng, Jinwei Gu, Qinsheng Zhang, Hanzi Mao 외 arxiv

Training generalist policies for robotic manipulation has shown great promise, as they enable language-conditioned, multi-task behaviors across diverse scenarios. However, evaluating these policies remains difficult beca…

Video Generation

Evaluating Real-World Robot Manipulation Policies in Simulation

2024-05-09 · Xuanlin Li, Kyle Hsu, Jiayuan Gu, Karl Pertsch 외

The field of robotics has made significant advances towards generalist robot manipulation policies. However, real-world evaluation of such policies is not scalable and faces reproducibility challenges, which are likely t…

Robotic GraspingRobot ManipulationRobot Manipulation Generalization

Benchmarking Vision-Language-Action Models on SO-101: Failure and Recovery Analysis

2026-06-07 · Yi Yu, Xinchuan Qiu arxiv

Vision-Language-Action (VLA) models have demonstrated strong generalization in robotic manipulation, yet existing evaluations are primarily conducted in simulation or on expensive robotic platforms, leaving their robustn…

Composable Deep Reinforcement Learning for Robotic Manipulation

2018-03-19 · Tuomas Haarnoja, Vitchyr Pong, Aurick Zhou, Murtaza Dalal 외

Model-free deep reinforcement learning has been shown to exhibit good performance in domains ranging from video games to simulated robotic manipulation and locomotion. However, model-free methods are known to perform poo…

Deep Reinforcement LearningQ-Learningreinforcement-learningReinforcement Learning+1