paper-with-me

홈 › Papers

Beyond Segmentation: Road Network Generation with Multi-Modal LLMs

2023-10-15 · Sumedh Rasal, Sanjay Kumar Boddhu

This paper introduces an innovative approach to road network generation through the utilization of a multi-modal Large Language Model (LLM). Our model is specifically designed to process aerial images of road layouts and produce detailed, navigable road networks within the input images. The core innovation of our system lies in the unique training methodology employed for the large language model to generate road networks as its output. This approach draws inspiration from the BLIP-2 architecture arXiv:2301.12597, leveraging pre-trained frozen image encoders and large language models to create a versatile multi-modal LLM. Our work also offers an alternative to the reasoning segmentation method proposed in the LISA paper arXiv:2308.00692. By training the large language model with our approach, the necessity for generating binary segmentation masks, as suggested in the LISA paper arXiv:2308.00692, is effectively eliminated. Experimental results underscore the efficacy of our multi-modal LLM in providing precise and valuable navigational guidance. This research represents a significant stride in bolstering autonomous navigation systems, especially in road network scenarios, where accurate guidance is of paramount importance.

📄 PDF Abstract BibTeX arXiv:2310.09755

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous NavigationLanguage ModelingLanguage ModellingLarge Language ModelReasoning Segmentation

Similar Papers 제목 키워드 기반

Adaptive-Mask Fusion Network for Segmentation of Drivable Road and Negative Obstacle With Untrustworthy Features

2023-04-27 · Zhen Feng, Yuchao Feng, Yanning Guo, Yuxiang Sun

Segmentation of drivable roads and negative obstacles is critical to the safe driving of autonomous vehicles. Currently, many multi-modal fusion methods have been proposed to improve segmentation accuracy, such as fusing…

Autonomous VehiclesSegmentation

LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

2023-11-01 · Wei-Ge Chen, Irina Spiridonova, Jianwei Yang, Jianfeng Gao 외

LLaVA-Interactive is a research prototype for multimodal human-AI interaction. The system can have multi-turn dialogues with human users by taking multimodal user inputs and generating multimodal responses. Importantly, …

AllImage GenerationImage SegmentationSemantic Segmentation

CORP: A Multi-Modal Dataset for Campus-Oriented Roadside Perception Tasks

2024-04-04 · Beibei Wang, Shuang Meng, Lu Zhang, Chenjie Wang 외

Numerous roadside perception datasets have been introduced to propel advancements in autonomous driving and intelligent transportation systems research and development. However, it has been observed that the majority of …

Autonomous DrivingInstance SegmentationSemantic Segmentation

Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation

2026-08-03 · Jiazhen Liu, Mingkuan Feng, Long Chen hf

MLLM-based segmentation faces a core segmentation trilemma: high segmentation performance, preserved dialogue ability, and fast inference. Embedding-prediction methods may disrupt language modeling through pixel-level ob…

Instruction Following

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

2026-05-12 · Haiwen Diao, Penghao Wu, Hanming Deng, Jiahao Wang 외 arxiv

Recent large vision-language models (VLMs) remain fundamentally constrained by a persistent dichotomy: understanding and generation are treated as distinct problems, leading to fragmented architectures, cascaded pipeline…