paper-with-me

홈 › Papers

NuScenes-MQA: Integrated Evaluation of Captions and QA for Autonomous Driving Datasets using Markup Annotations

2023-12-11 · Yuichi Inoue, Yuki Yada, Kotaro Tanahashi, Yu Yamaguchi

Visual Question Answering (VQA) is one of the most important tasks in autonomous driving, which requires accurate recognition and complex situation evaluations. However, datasets annotated in a QA format, which guarantees precise language generation and scene recognition from driving scenes, have not been established yet. In this work, we introduce Markup-QA, a novel dataset annotation technique in which QAs are enclosed within markups. This approach facilitates the simultaneous evaluation of a model's capabilities in sentence generation and VQA. Moreover, using this annotation methodology, we designed the NuScenes-MQA dataset. This dataset empowers the development of vision language models, especially for autonomous driving tasks, by focusing on both descriptive capabilities and precise QA. The dataset is available at https://github.com/turingmotors/NuScenes-MQA.

📄 PDF Abstract BibTeX arXiv:2312.06352

Code (1)

turingmotors/nuscenes-mqa 공식 구현

Tasks

Autonomous DrivingDescriptiveQuestion AnsweringScene RecognitionSentenceText GenerationVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

On Offline Evaluation of 3D Object Detection for Autonomous Driving

2023-08-24 · Tim Schreier, Katrin Renz, Andreas Geiger, Kashyap Chitta

Prior work in 3D object detection evaluates models using offline metrics like average precision since closed-loop online evaluation on the downstream driving task is costly. However, it is unclear how indicative offline …

3D Object DetectionAutonomous DrivingObjectobject-detection+1

A Low-Rank Method for Vision Language Model Hallucination Mitigation in Autonomous Driving

2025-11-09 · Keke Long, Jiacheng Guo, Tianyun Zhang, Hongkai Yu 외 arxiv

Vision Language Models (VLMs) are increasingly used in autonomous driving to help understand traffic scenes, but they sometimes produce hallucinations, which are false details not grounded in the visual input. Detecting …

Autonomous Driving

Rethinking the Open-Loop Evaluation of End-to-End Autonomous Driving in nuScenes

2023-05-17 · Jiang-Tian Zhai, Ze Feng, Jinhao Du, Yongqiang Mao 외

Modern autonomous driving systems are typically divided into three main tasks: perception, prediction, and planning. The planning task involves predicting the trajectory of the ego vehicle based on inputs from both inter…

Autonomous DrivingTrajectory Planning

NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving

2025-04-04 · Kexin Tian, Jingrui Mao, Yunlong Zhang, Jiwan Jiang 외

Recent advancements in Vision-Language Models (VLMs) have demonstrated strong potential for autonomous driving tasks. However, their spatial understanding and reasoning-key capabilities for autonomous driving-still exhib…

3d scene graph generationAutonomous DrivingGraph GenerationScene Graph Generation+1

nuScenes Revisited: Progress and Challenges in Autonomous Driving

2025-12-02 · Whye Kit Fong, Venice Erin Liong, Kok Seang Tan, Holger Caesar arxiv

Autonomous Vehicles (AV) and Advanced Driver Assistance Systems (ADAS) have been revolutionized by Deep Learning. As a data-driven approach, Deep Learning relies on vast amounts of driving data, typically labeled in grea…

Autonomous VehiclesAutonomous Driving