paper-with-me

홈 › Papers

ChatBEV: A Visual Language Model that Understands BEV Maps

2025-03-18 · Qingyao Xu, Siheng Chen, Guang Chen, Yanfeng Wang, Ya zhang

Traffic scene understanding is essential for intelligent transportation systems and autonomous driving, ensuring safe and efficient vehicle operation. While recent advancements in VLMs have shown promise for holistic scene understanding, the application of VLMs to traffic scenarios, particularly using BEV maps, remains under explored. Existing methods often suffer from limited task design and narrow data amount, hindering comprehensive scene understanding. To address these challenges, we introduce ChatBEV-QA, a novel BEV VQA benchmark contains over 137k questions, designed to encompass a wide range of scene understanding tasks, including global scene understanding, vehicle-lane interactions, and vehicle-vehicle interactions. This benchmark is constructed using an novel data collection pipeline that generates scalable and informative VQA data for BEV maps. We further fine-tune a specialized vision-language model ChatBEV, enabling it to interpret diverse question prompts and extract relevant context-aware information from BEV maps. Additionally, we propose a language-driven traffic scene generation pipeline, where ChatBEV facilitates map understanding and text-aligned navigation guidance, significantly enhancing the generation of realistic and consistent traffic scenarios. The dataset, code and the fine-tuned model will be released.

📄 PDF Abstract BibTeX arXiv:2503.13938

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingLanguage ModelingLanguage ModellingScene GenerationScene UnderstandingVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

CLIP also Understands Text: Prompting CLIP for Phrase Understanding

2022-10-11 · An Yan, Jiacheng Li, Wanrong Zhu, Yujie Lu 외

Contrastive Language-Image Pretraining (CLIP) efficiently learns visual concepts by pre-training with natural language supervision. CLIP and its visual encoder have been explored on various vision and language tasks and …

ClusteringTransfer Learning

Learning Sim-to-Real Dense Object Descriptors for Robotic Manipulation

2023-04-18 · Hoang-Giang Cao, Weihao Zeng, I-Chen Wu

It is crucial to address the following issues for ubiquitous robotics manipulation applications: (a) vision-based manipulation tasks require the robot to visually learn and understand the object with rich information lik…

Object

Multimodal LLMs Struggle with Basic Visual Network Analysis: a VNA Benchmark

2024-05-10 · Evan M. Williams, Kathleen M. Carley

We evaluate the zero-shot ability of GPT-4 and LLaVa to perform simple Visual Network Analysis (VNA) tasks on small-scale graphs. We evaluate the Vision Language Models (VLMs) on 5 tasks related to three foundational net…

Multimodal Hierarchical Reinforcement Learning Policy for Task-Oriented Visual Dialog

2018-05-08 · WS 2018 7 · Jiaping Zhang, Tiancheng Zhao, Zhou Yu

Creating an intelligent conversational system that understands vision and language is one of the ultimate goals in Artificial Intelligence (AI)~\cite{winograd1972understanding}. Extensive research has focused on vision-t…

Hierarchical Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)+2

Aligning Brain Signals with Multimodal Speech and Vision Embeddings

2025-10-29 · Kateryna Shapovalenko, Quentin Auster arxiv

When we hear the word "house", we don't just process sound, we imagine walls, doors, memories. The brain builds meaning through layers, moving from raw acoustics to rich, multimodal associations. Inspired by this, we bui…