paper-with-me

Papers

RobotArena $\infty$: Scalable Robot Benchmarking via Real-to-Sim Translation

2025-10-27 · Yash Jangir, Yidi Zhang, Pang-Chi Lo, Kashu Yamazaki, Chenyu Zhang, Kuan-Hsun Tu, Tsung-Wei Ke, Lei Ke, Yonatan Bisk, Katerina Fragkiadaki arxiv

The pursuit of robot generalists, agents capable of performing diverse tasks across diverse environments, demands rigorous and scalable evaluation. Yet real-world testing of robot policies remains fundamentally constrained: it is labor-intensive, slow, unsafe at scale, and difficult to reproduce. As policies expand in scope and complexity, these barriers only intensify, since defining "success" in robotics often hinges on nuanced human judgments of execution quality. We introduce RobotArena Infinity, a new benchmarking framework that overcomes these challenges by shifting vision-language-action (VLA) evaluation into large-scale simulated environments augmented with online human feedback. Leveraging advances in vision-language models, 2D-to-3D generative modeling, and differentiable rendering, our approach automatically converts video demonstrations from widely used robot datasets into simulated counterparts. Within these digital twins, we assess VLA policies using both automated vision-language-model-guided scoring and scalable human preference judgments collected from crowdworkers, transforming human involvement from tedious scene setup, resetting, and safety supervision into lightweight preference comparisons. To measure robustness, we systematically perturb simulated environments along multiple axes, including textures and object placements, stress-testing policy generalization under controlled variation. The result is a continuously evolving, reproducible, and scalable benchmark for real-world-trained robot manipulation policies, addressing a critical missing capability in today's robotics landscape.

📄 PDF Abstract BibTeX arXiv:2510.23571

Code (0)

등록된 구현이 없습니다.

Tasks

Robot Manipulation

Similar Papers 제목 키워드 기반

PolaRiS: Scalable Real-to-Sim Evaluations for Generalist Robot Policies

2025-12-18 · Arhan Jain, Mingtong Zhang, Kanav Arora, William Chen 외 arxiv

A significant challenge for robot learning research is our ability to accurately measure and compare the performance of robot policies. Benchmarking in robotics is historically challenging due to the stochasticity, repro…

RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies

2026-04-10 · Xuning Yang, Rishit Dagli, Alex Zook, Hugo Hadfield 외 arxiv

The pursuit of general-purpose robotics has yielded impressive foundation models, yet simulation-based benchmarking remains a bottleneck due to rapid performance saturation and a lack of true generalization testing. Exis…

ManipulationNet: An Infrastructure for Benchmarking Real-World Robot Manipulation with Physical Skill Challenges and Embodied Multimodal Reasoning

2026-03-04 · Yiting Chen, Kenneth Kimble, Edward H. Adelson, Tamim Asfour 외 arxiv

Dexterous manipulation enables robots to purposefully alter the physical world, transforming them from passive observers into active agents in unstructured environments. This capability is the cornerstone of physical art…

Multimodal ReasoningRobot Manipulation

RSLCPP -- Deterministic Simulations Using ROS 2

2026-01-11 · Simon Sagmeister, Marcel Weinmann, Phillip Pitschi, Markus Lienkamp arxiv

Simulation is crucial in real-world robotics, offering safe, scalable, and efficient environments for developing a variety of robotic applications. While the Robot Operating System (ROS) has been widely adopted as the ba…

SPROUT: Self-Progressing Robust Training

2019-09-25 · Minhao Cheng, Pin-Yu Chen, Sijia Liu, Shiyu Chang 외

Enhancing model robustness under new and even adversarial environments is a crucial milestone toward building trustworthy and reliable machine learning systems. Current robust training methods such as adversarial trainin…

Adversarial Robustness