paper-with-me

홈 › Papers

Driving with InternVL: Oustanding Champion in the Track on Driving with Language of the Autonomous Grand Challenge at CVPR 2024

2024-12-10 · Jiahan Li, Zhiqi Li, Tong Lu

This technical report describes the methods we employed for the Driving with Language track of the CVPR 2024 Autonomous Grand Challenge. We utilized a powerful open-source multimodal model, InternVL-1.5, and conducted a full-parameter fine-tuning on the competition dataset, DriveLM-nuScenes. To effectively handle the multi-view images of nuScenes and seamlessly inherit InternVL's outstanding multimodal understanding capabilities, we formatted and concatenated the multi-view images in a specific manner. This ensured that the final model could meet the specific requirements of the competition task while leveraging InternVL's powerful image understanding capabilities. Meanwhile, we designed a simple automatic annotation strategy that converts the center points of objects in DriveLM-nuScenes into corresponding bounding boxes. As a result, our single model achieved a score of 0.6002 on the final leadboard.

📄 PDF Abstract BibTeX arXiv:2412.07247

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DINO-SD: Champion Solution for ICRA 2024 RoboDepth Challenge

2024-05-27 · Yifan Mao, Ming Li, Jian Liu, Jiayang Liu 외

Surround-view depth estimation is a crucial task aims to acquire the depth maps of the surrounding views. It has many applications in real world scenarios such as autonomous driving, AR/VR and 3D reconstruction, etc. How…

3D ReconstructionAutonomous DrivingDepth Estimation

Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

2024-10-21 · Zhangwei Gao, Zhe Chen, Erfei Cui, Yiming Ren 외

Multimodal large language models (MLLMs) have demonstrated impressive performance in vision-language tasks across a broad spectrum of domains. However, the large model scale and associated high computational costs pose s…

Autonomous Driving

Unlock the Power of Unlabeled Data in Language Driving Model

2025-03-13 · Chaoqun Wang, Jie Yang, Xiaobin Hong, Ruimao Zhang

Recent Vision-based Large Language Models~(VisionLLMs) for autonomous driving have seen rapid advancements. However, such promotion is extremely dependent on large-scale high-quality annotated data, which is costly and l…

Autonomous DrivingQuestion Answering

Cross-Stage Coherence in Hierarchical Driving VQA: Explicit Baselines and Learned Gated Context Projectors

2026-04-24 · Gautam Kumar Jain, Carsten Markgraf, Julian Stähler arxiv

Graph Visual Question Answering (GVQA) for autonomous driving organizes reasoning into ordered stages, namely Perception, Prediction, and Planning, where planning decisions should remain consistent with the model's own p…

Visual Question AnsweringAutonomous DrivingDomain Adaptation

Precise Drive with VLM: First Prize Solution for PRCV 2024 Drive LM challenge

2024-11-05 · Bin Huang, Siyu Wang, Yuanpeng Chen, Yidan Wu 외

This technical report outlines the methodologies we applied for the PRCV Challenge, focusing on cognition and decision-making in driving scenarios. We employed InternVL-2.0, a pioneering open-source multi-modal model, an…

Autonomous DrivingDecision Making