paper-with-me

Papers

Embodied Scene Understanding for Vision Language Models via MetaVQA

2025-01-15 · CVPR 2025 1 · Weizhen Wang, Chenda Duan, Zhenghao Peng, Yuxin Liu, Bolei Zhou

Vision Language Models (VLMs) demonstrate significant potential as embodied AI agents for various mobility applications. However, a standardized, closed-loop benchmark for evaluating their spatial reasoning and sequential decision-making capabilities is lacking. To address this, we present MetaVQA: a comprehensive benchmark designed to assess and enhance VLMs' understanding of spatial relationships and scene dynamics through Visual Question Answering (VQA) and closed-loop simulations. MetaVQA leverages Set-of-Mark prompting and top-down view ground-truth annotations from nuScenes and Waymo datasets to automatically generate extensive question-answer pairs based on diverse real-world traffic scenarios, ensuring object-centric and context-rich instructions. Our experiments show that fine-tuning VLMs with the MetaVQA dataset significantly improves their spatial reasoning and embodied scene comprehension in safety-critical simulations, evident not only in improved VQA accuracies but also in emerging safety-aware driving maneuvers. In addition, the learning demonstrates strong transferability from simulation to real-world observation. Code and data will be publicly available at https://metadriverse.github.io/metavqa .

📄 PDF Abstract BibTeX arXiv:2501.09167

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingQuestion AnsweringScene UnderstandingSequential Decision MakingSpatial ReasoningVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models

2024-06-09 · Mengfei Du, Binhao Wu, Zejun Li, Xuanjing Huang 외

The recent rapid development of Large Vision-Language Models (LVLMs) has indicated their potential for embodied tasks.However, the critical skill of spatial understanding in embodied environments has not been thoroughly …

Benchmarking

Perception Matters: Detecting Perception Failures of VQA Models Using Metamorphic Testing

2021-06-19 · CVPR 2021 1 · Yuanyuan Yuan, Shuai Wang, Mingyue Jiang, Tsong Yueh Chen

Visual question answering (VQA) takes an image and a natural-language question as input and returns a natural-language answer. To date, VQA models are primarily assessed by their accuracy on high-level reasoning ques…

BenchmarkingDNN TestingQuestion AnsweringVisual Question Answering+1

Embodied Understanding of Driving Scenarios

2024-03-07 · Yunsong Zhou, Linyan Huang, Qingwen Bu, Jia Zeng 외

Embodied scene understanding serves as the cornerstone for autonomous agents to perceive, interpret, and respond to open driving scenarios. Such understanding is typically founded upon Vision-Language Models (VLMs). Neve…

Autonomous DrivingLanguage ModelingLanguage ModellingScene Understanding

HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding

2025-03-17 · Jiahe Zhao, Ruibing Hou, Zejie Tian, Hong Chang 외

We propose a new task to benchmark human-in-scene understanding for embodied agents: Human-In-Scene Question Answering (HIS-QA). Given a human motion within a 3D scene, HIS-QA requires the agent to comprehend human state…

Question AnsweringScene Understanding

YouRefIt: Embodied Reference Understanding with Language and Gesture

2021-09-08 · ICCV 2021 10 · Yixin Chen, Qing Li, Deqian Kong, Yik Lun Kei 외

We study the understanding of embodied reference: One agent uses both language and gesture to refer to an object to another agent in a shared physical environment. Of note, this new visual task requires understanding mul…