paper-with-me

홈 › Papers

ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation

2025-05-14 · Enyu Zhao, Vedant Raval, Hejia Zhang, Jiageng Mao, Zeyu Shangguan, Stefanos Nikolaidis, Yue Wang, Daniel Seita

Vision-Language Models (VLMs) have revolutionized artificial intelligence and robotics due to their commonsense reasoning capabilities. In robotic manipulation, VLMs are used primarily as high-level planners, but recent work has also studied their lower-level reasoning ability, which refers to making decisions about precise robot movements. However, the community currently lacks a clear and common benchmark that can evaluate how well VLMs can aid low-level reasoning in robotics. Consequently, we propose a novel benchmark, ManipBench, to evaluate the low-level robot manipulation reasoning capabilities of VLMs across various dimensions, including how well they understand object-object interactions and deformable object manipulation. We extensively test 33 representative VLMs across 10 model families on our benchmark, including variants to test different model sizes. Our evaluation shows that the performance of VLMs significantly varies across tasks, and there is a strong correlation between this performance and trends in our real-world manipulation tasks. It also shows that there remains a significant gap between these models and human-level understanding. See our website at: https://manipbench.github.io.

📄 PDF Abstract BibTeX arXiv:2505.09698

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingDeformable Object ManipulationObjectRobot Manipulation

Similar Papers 제목 키워드 기반

ManipShield: A Unified Framework for Image Manipulation Detection, Localization and Explanation

2025-11-18 · Zitong Xu, Huiyu Duan, Xiaoyu Wang, Zhaolin Cai 외 arxiv

With the rapid advancement of generative models, powerful image editing methods now enable diverse and highly realistic image manipulations that far surpass traditional deepfake techniques, posing new challenges for mani…

Image Manipulation DetectionImage Editing

A Benchmarking Study of Vision-based Robotic Grasping Algorithms

2025-03-14 · Bharath K Rameshbabu, Sumukh S Balakrishna, Brian Flynn, Vinarak Kapoor 외

We present a benchmarking study of vision-based robotic grasping algorithms with distinct approaches, and provide a comparative analysis. In particular, we compare two machine-learning-based and two analytical algorithms…

BenchmarkingRobotic Grasping

RobotArena $\infty$: Scalable Robot Benchmarking via Real-to-Sim Translation

2025-10-27 · Yash Jangir, Yidi Zhang, Pang-Chi Lo, Kashu Yamazaki 외 arxiv

The pursuit of robot generalists, agents capable of performing diverse tasks across diverse environments, demands rigorous and scalable evaluation. Yet real-world testing of robot policies remains fundamentally constrain…

Robot Manipulation

Mind and Motion Aligned: A Joint Evaluation IsaacSim Benchmark for Task Planning and Low-Level Policies in Mobile Manipulation

2025-08-21 · Nikita Kachaev, Andrei Spiridonov, Andrey Gorodetsky, Kirill Muravyev 외 arxiv

Benchmarks are crucial for evaluating progress in robotics and embodied AI. However, a significant gap exists between benchmarks designed for high-level language instruction following, which often assume perfect low-leve…

Instruction Following

ManiSoft: Towards Vision-Language Manipulation for Soft Continuum Robotics

2026-05-18 · Ziyu Wei, Luting Wang, Chen Gao, Li Wen 외 arxiv

Most existing vision-language manipulation research targets rigid robotic arms, whose fixed morphology limits adaptability in cluttered or confined spaces. Soft robotic arms offer an appealing alternative due to their de…

Reinforcement Learning