paper-with-me

홈 › Papers

HiLM-D: Towards High-Resolution Understanding in Multimodal Large Language Models for Autonomous Driving

2023-09-11 · Xinpeng Ding, Jianhua Han, Hang Xu, Wei zhang, Xiaomeng Li

Autonomous driving systems generally employ separate models for different tasks resulting in intricate designs. For the first time, we leverage singular multimodal large language models (MLLMs) to consolidate multiple autonomous driving tasks from videos, i.e., the Risk Object Localization and Intention and Suggestion Prediction (ROLISP) task. ROLISP uses natural language to simultaneously identify and interpret risk objects, understand ego-vehicle intentions, and provide motion suggestions, eliminating the necessity for task-specific architectures. However, lacking high-resolution (HR) information, existing MLLMs often miss small objects (e.g., traffic cones) and overly focus on salient ones (e.g., large trucks) when applied to ROLISP. We propose HiLM-D (Towards High-Resolution Understanding in MLLMs for Autonomous Driving), an efficient method to incorporate HR information into MLLMs for the ROLISP task. Especially, HiLM-D integrates two branches: (i) the low-resolution reasoning branch, can be any MLLMs, processes low-resolution videos to caption risk objects and discern ego-vehicle intentions/suggestions; (ii) the high-resolution perception branch (HR-PB), prominent to HiLM-D,, ingests HR images to enhance detection by capturing vision-specific HR feature maps and prioritizing all potential risks over merely salient objects. Our HR-PB serves as a plug-and-play module, seamlessly fitting into current MLLMs. Experiments on the ROLISP benchmark reveal HiLM-D's notable advantage over leading MLLMs, with improvements of 4.8% in BLEU-4 for captioning and 17.2% in mIoU for detection.

📄 PDF Abstract BibTeX arXiv:2309.05186

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingObject Localization

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation

2025-09-23 · Shufan Li, Jiuxiang Gu, Kangning Liu, Zhe Lin 외 arxiv

We propose Lavida-O, a unified Masked Diffusion Model (MDM) for multimodal understanding and generation. Unlike existing multimodal MDMs such as MMaDa and Muddit which only support simple image-level understanding tasks …

Text-to-Image GenerationMultimodal ReasoningImage Editing

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding

2025-01-14 · Zhaokai Wang, Xizhou Zhu, Xue Yang, Gen Luo 외

Image pyramids are widely adopted in top-performing methods to obtain multi-scale features for precise visual perception and understanding. However, current image pyramids use the same large-scale model to process multip…

image-classificationImage ClassificationLarge Language ModelMultimodal Large Language Model+2

Does Table Source Matter? Benchmarking and Improving Multimodal Scientific Table Understanding and Reasoning

2025-01-22 · Bohao Yang, Yingji Zhang, Dong Liu, André Freitas 외

Recent large language models (LLMs) have advanced table understanding capabilities but rely on converting tables into text sequences. While multimodal large language models (MLLMs) enable direct visual processing, they f…

Benchmarking

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

2024-04-25 · Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye 외

In this report, we introduce InternVL 1.5, an open-source multimodal large language model (MLLM) to bridge the capability gap between open-source and proprietary commercial models in multimodal understanding. We introduc…

4kLanguage ModelingLanguage ModellingLarge Language Model+4

OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation

2025-09-03 · Han Li, Xinyu Peng, Yaoming Wang, Zelin Peng 외 arxiv

We introduce OneCAT, a unified multimodal model that seamlessly integrates understanding, generation, and editing within a novel, pure decoder-only transformer architecture. Our framework uniquely eliminates the need for…

multimodal generation