paper-with-me

홈 › Papers

MAPLM: A Real-World Large-Scale Vision-Language Benchmark for Map and Traffic Scene Understanding

2024-01-01 · CVPR 2024 1 · Xu Cao, Tong Zhou, Yunsheng Ma, Wenqian Ye, Can Cui, Kun Tang, Zhipeng Cao, Kaizhao Liang, Ziran Wang, James M. Rehg, Chao Zheng

Vision-language generative AI has demonstrated remarkable promise for empowering cross-modal scene understanding of autonomous driving and high-definition (HD) map systems. However current benchmark datasets lack multi-modal point cloud image and language data pairs. Recent approaches utilize visual instruction learning and cross-modal prompt engineering to expand vision-language models into this domain. In this paper we propose a new vision-language benchmark that can be used to finetune traffic and HD map domain-specific foundation models. Specifically we annotate and leverage large-scale broad-coverage traffic and map data extracted from huge HD map annotations and use CLIP and LLaMA-2 / Vicuna to finetune a baseline model with instruction-following data. Our experimental results across various algorithms reveal that while visual instruction-tuning large language models (LLMs) can effectively learn meaningful representations from MAPLM-QA there remains significant room for further advancements. To facilitate applying LLMs and multi-modal data into self-driving research we will release our visual-language QA data and the baseline models at GitHub.com/LLVM-AD/MAPLM.

📄 PDF Abstract BibTeX

Code (1)

llvm-ad/maplm 공식 구현

Tasks

Autonomous DrivingInstruction FollowingPrompt EngineeringScene Understanding

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Boosting Zero-shot Stereo Matching using Large-scale Mixed Images Sources in the Real World

2025-05-13 · Yuran Wang, Yingping Liang, Ying Fu

Stereo matching methods rely on dense pixel-wise ground truth labels, which are laborious to obtain, especially for real-world datasets. The scarcity of labeled data and domain gaps between synthetic and real-world image…

Depth EstimationMonocular Depth EstimationStereo Matching

MoWorld: A Flash World Model

2026-07-07 · Team Moxin, Deyi Ji, Tianrun Chen, Xin Zhang 외 arxiv

The future of World Models depends not only on scaling model capability, but also on scaling practicality and inference efficiency. High-frame-rate inference enables responsive perception, planning, and control in real-w…

The ParallelEye Dataset: Constructing Large-Scale Artificial Scenes for Traffic Vision Research

2017-12-22 · Xuan Li, Kunfeng Wang, Yonglin Tian, Lan Yan 외

Video image datasets are playing an essential role in design and evaluation of traffic vision algorithms. Nevertheless, a longstanding inconvenience concerning image datasets is that manually collecting and annotating la…

Instance SegmentationObject TrackingOptical Flow EstimationSemantic Segmentation

Towards Real-World Ultrasound Understanding: Large Vision-Language Models from Multi-Image Examinations with Long-Form Reports

2026-07-02 · Bingcong Yan, Chunlei Li, Jingliang Hu, Yilei Shi 외 arxiv

Large vision-language models (LVLMs) have achieved strong performance across many medical imaging tasks, yet their application to ultrasound remains limited due to its inherent complexity and variability. In this work, w…

ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining

2026-06-15 · Hao Li, Ganlong Zhao, Yufei Liu, Haotian Hou 외 arxiv

Vision-Language-Action (VLA) models benefit from large-scale and diverse embodied data, yet scaling robot trajectory collection is costly and labor-intensive. Recent advances show that large-scale egocentric human videos…