paper-with-me

Papers

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation

2026-06-11 · Dayu Xia, Yue Shi, Yao Mu, Huiting Ji, Chaofan Ma, Yingjie Zhou, Hua Chen, Yang Liu, Jiezhang Cao, Guangtao Zhai arxiv

Vision-language models (VLMs) are increasingly explored as visual critics, reward generators, and failure detectors in robotic manipulation. These roles implicitly require models to judge not only final task success, but also how a manipulation execution is physically and temporally progressing. However, existing evaluations fail to test whether VLMs possess fine-grained process understanding. To address this gap, we present RoboProcessBench, a benchmark for process-aware understanding in vision-language robotic manipulation. RoboProcessBench decomposes such capability into two complementary dimensions, \emph{static monitoring} and \emph{dynamic reasoning}, instantiated as 12 diagnostic question families covering phase, contact, motion, coordination, primitive-local progress, temporal order, outcome, and primitive-level transitions. Built from physically grounded execution traces, the curated benchmark corpus ProcessData contains \textasciitilde 58k question-answer pairs across 260 manipulation tasks, which is further split into ProcessData-SFT and ProcessData-Eval for post-training and evaluation purposes. Extensive evaluation of various VLMs on ProcessData-Eval reveals broad limitations across 12 diagnostic task families, suggesting current models still lack robust process-aware understanding of manipulation executions. But with ProcessData-SFT, the post-trained \textit{Qwen2.5-VL-7B} and \textit{InternVL-3-8B} exhibit consistent gains on local state, motion, progress, and primitive-aware cues. These results demonstrate that RoboProcessBench serves as both an evaluation benchmark and a learnable supervision source for developing VLMs capable of monitoring and evaluating robotic manipulation processes. Project webpage: \href{https://processbench-2026.github.io/RoboProcessBench-Web/}{https://processbench-2026.github.io}.

📄 PDF Abstract BibTeX arXiv:2606.13040

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Benchmarking Vision Language Models for Cultural Understanding

2024-07-15 · Shravan Nayak, Kanishk Jain, Rabiul Awal, Siva Reddy 외

Foundation models and vision-language pre-training have notably advanced Vision Language Models (VLMs), enabling multimodal processing of visual and linguistic data. However, their performance has been typically assessed…

BenchmarkingQuestion AnsweringScene UnderstandingVisual Question Answering

Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding

2025-09-26 · Vahid Mirjalili, Ramin Giahi, Sriram Kollipara, Akshay Kekuda 외 arxiv

Spatial understanding is a critical capability for vision foundation models. While recent advances in large vision models or vision-language models (VLMs) have expanded recognition capabilities, most benchmarks emphasize…

Relational ReasoningScene UnderstandingSpatial Reasoning

Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning

2024-12-18 · Yingjie Zhu, Xuefeng Bai, Kehai Chen, Yang Xiang 외

Large Vision-Language Models (LVLMs) have demonstrated remarkable performance across diverse tasks. Despite great success, recent studies show that LVLMs encounter substantial limitations when engaging with visual graphs…

BenchmarkingGraph LearningSelf-Supervised Learning

SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision

2025-08-05 · Zhaoxu Li, Chenqi Kong, Yi Yu, Qiangqiang Wu 외 arxiv

Large Vision-Language Models (LVLMs) recently achieve significant breakthroughs in understanding complex visual-textual contexts. However, hallucination issues still limit their real-world applicability. Although previou…

Scene Understanding

ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness

2026-07-26 · Qiao Yan, Yihan Wang, Zhenghao Xing, Jiaqi Xu 외 arxiv

Autonomous driving under adverse weather remains a critical challenge, yet existing vision-language benchmarks mainly evaluate under standard conditions, synthetic corruptions, or single modality. As a result, it remains…

Reinforcement LearningScene UnderstandingAutonomous Driving