paper-with-me

홈 › Papers

RoboVista: Evaluating Vision Language Models for Diverse Robot Applications

2026-07-06 · Shuangyu Xie, Kaiyuan Chen, Ziyang Chen, Simeon Adebola, Yixuan Huang, Zehan Ma, Tianshuang Qiu, Wentao Yuan, Dhruv Shah, Pannag R. Sanketi, Ken Goldberg arxiv

Diverse applications for robotics, such as industry and agriculture, require robots to operate across various embodiments, changing visual conditions, and complex planning. Vision-Language Models (VLMs) offer a promising foundation for general-purpose and interpretable robotic reasoning. Aligning VLMs with diverse robot applications requires a modular understanding of the individual decision components that underlie robotic behavior. Capturing such structure is challenging for conventional robot benchmarks that are primarily based on teleoperated, end-to-end datasets. We propose Robot Question Answering (RQA), a modular evaluation framework and RoboVista, a benchmark curated from real robotic systems, research papers, and expert annotations. RoboVista contains 474 Visual Question Answering (VQA) instances with human annotated reasoning and covers 39 unique task types in agricultural, industrial, domestic, surgical robotics, autonomous driving, and open robot datasets. Experiments on RoboVista show that state-of-the-art VLMs exhibit substantial gaps. Physical robot experiments suggest strong correlation between RoboVista performance and real-world task execution.

📄 PDF Abstract BibTeX arXiv:2607.04610

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringAutonomous Driving

Similar Papers 제목 키워드 기반

LADEV: A Language-Driven Testing and Evaluation Platform for Vision-Language-Action Models in Robotic Manipulation

2024-10-07 · Zhijie Wang, Zhehua Zhou, Jiayang Song, Yuheng Huang 외

Building on the advancements of Large Language Models (LLMs) and Vision Language Models (VLMs), recent research has introduced Vision-Language-Action (VLA) models as an integrated solution for robotic manipulation tasks.…

Vision-Language-Action

Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks

2024-11-04 · Pranav Guruprasad, Harshvardhan Sikka, Jaewoo Song, Yangyue Wang 외

Vision-language-action (VLA) models represent a promising direction for developing general-purpose robotic systems, demonstrating the ability to combine visual understanding, language comprehension, and action generation…

Action GenerationBenchmarkingPrompt EngineeringVision-Language-Action

Colosseum V2: Benchmarking Generalization for Vision Language Action Models

2026-05-26 · Jeremy Morgan, Prajwal Vijay, Hyeonho Oh, Jincen Song 외 arxiv

Vision-Language-Action (VLA) models demonstrate promising generalization in robotic manipulation, driven by advances in large-scale vision and language pre-training. This progress can be misleading. Despite the zero-shot…

Embodied Red Teaming for Auditing Robotic Foundation Models

2024-11-27 · Sathwik Karnik, Zhang-Wei Hong, Nishant Abhangi, Yen-Chen Lin 외

Language-conditioned robot models have the potential to enable robots to perform a wide range of tasks based on natural language instructions. However, assessing their safety and effectiveness remains challenging because…

Red Teaming

Emergence of Human to Robot Transfer in Vision-Language-Action Models

2025-12-27 · Simar Kareer, Karl Pertsch, James Darpinian, Judy Hoffman 외 arxiv

Vision-language-action (VLA) models can enable broad open world generalization, but require large and diverse datasets. It is appealing to consider whether some of this data can come from human videos, which cover divers…