paper-with-me

Papers

AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World

2025-03-31 · Zhiyuan Zhou, Pranav Atreya, You Liang Tan, Karl Pertsch, Sergey Levine

Scalable and reproducible policy evaluation has been a long-standing challenge in robot learning. Evaluations are critical to assess progress and build better policies, but evaluation in the real world, especially at a scale that would provide statistically reliable results, is costly in terms of human time and hard to obtain. Evaluation of increasingly generalist robot policies requires an increasingly diverse repertoire of evaluation environments, making the evaluation bottleneck even more pronounced. To make real-world evaluation of robotic policies more practical, we propose AutoEval, a system to autonomously evaluate generalist robot policies around the clock with minimal human intervention. Users interact with AutoEval by submitting evaluation jobs to the AutoEval queue, much like how software jobs are submitted with a cluster scheduling system, and AutoEval will schedule the policies for evaluation within a framework supplying automatic success detection and automatic scene resets. We show that AutoEval can nearly fully eliminate human involvement in the evaluation process, permitting around the clock evaluations, and the evaluation results correspond closely to ground truth evaluations conducted by hand. To facilitate the evaluation of generalist policies in the robotics community, we provide public access to multiple AutoEval scenes in the popular BridgeData robot setup with WidowX robot arms. In the future, we hope that AutoEval scenes can be set up across institutions to form a diverse and distributed evaluation network.

📄 PDF Abstract BibTeX arXiv:2503.24278

Code (1)

zhouzypaul/auto_eval 공식 구현 jax

Tasks

Robot ManipulationScheduling

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Eval-Actions: Fine-Grained Execution Quality Evaluation for Robotic Manipulation

2026-01-26 · Mengyuan Liu, Juyi Sheng, Peiming Li, Ziyi Wang 외 arxiv

Although Vision--Action (VA) and Vision--Language--Action (VLA) policies have advanced robotic manipulation, their evaluation remains dominated by binary success rates, which obscure process-level differences among execu…

RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation

2023-06-20 · Konstantinos Bousmalis, Giulia Vezzani, Dushyant Rao, Coline Devin 외

The ability to leverage heterogeneous robotic experience from different robots and tasks to quickly master novel skills and embodiments has the potential to transform robot learning. Inspired by recent advances in founda…

Robot Policy Evaluation for Sim-to-Real Transfer: A Benchmarking Perspective

2025-08-14 · Xuning Yang, Clemens Eppner, Jonathan Tremblay, Dieter Fox 외 arxiv

Current vision-based robotics simulation benchmarks have significantly advanced robotic manipulation research. However, robotics is fundamentally a real-world problem, and evaluation for real-world applications has lagge…

GR-Dexter Technical Report

2025-12-30 · Ruoshi Wen, Guangzeng Chen, Zhongren Cui, Min Du 외 arxiv

Vision-language-action (VLA) models have enabled language-conditioned, long-horizon robot manipulation, but most existing systems are limited to grippers. Scaling VLA policies to bimanual robots with high degree-of-freed…

Robot Manipulation

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

2026-07-05 · Tianxing Chen, Yue Chen, Zixuan Li, Junyuan Tang 외 arxiv

Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their capabilities. Many rely on simple, short-horizon, or skill-narrow tasks with limited…

Instruction FollowingRobot Manipulation