Leveraging Multimodal-LLMs Assisted by Instance Segmentation for Intelligent Traffic Monitoring
A robust and efficient traffic monitoring system is essential for smart cities and Intelligent Transportation Systems (ITS), using sensors and cameras to track vehicle movements, optimize traffic flow, reduce congestion, enhance road safety, and enable real-time adaptive traffic control. Traffic monitoring models must comprehensively understand dynamic urban conditions and provide an intuitive user interface for effective management. This research leverages the LLaVA visual grounding multimodal large language model (LLM) for traffic monitoring tasks on the real-time Quanser Interactive Lab simulation platform, covering scenarios like intersections, congestion, and collisions. Cameras placed at multiple urban locations collect real-time images from the simulation, which are fed into the LLaVA model with queries for analysis. An instance segmentation model integrated into the cameras highlights key elements such as vehicles and pedestrians, enhancing training and throughput. The system achieves 84.3% accuracy in recognizing vehicle locations and 76.4% in determining steering direction, outperforming traditional models.
Code (0)
등록된 구현이 없습니다.
Tasks
Instance SegmentationLanguage ModelingLanguage ModellingLarge Language ModelMultimodal Large Language ModelSemantic SegmentationVisual GroundingSimilar Papers 제목 키워드 기반
MLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge
In this technical report, we present a training-free framework for audio-guided video object segmentation, which integrates Multimodal Large Language Models (MLLMs) with SAM-based segmentation models. We decompose the ta…
Video Object SegmentationMultimodal ReasoningVideo SegmentationAll in One: Visual-Description-Guided Unified Point Cloud Segmentation
Unified segmentation of 3D point clouds is crucial for scene understanding, but is hindered by its sparse structure, limited annotations, and the challenge of distinguishing fine-grained object classes in complex environ…
Point Cloud SegmentationPanoptic SegmentationScene UnderstandingPoint CloudsBoxSeg: Quality-Aware and Peer-Assisted Learning for Box-supervised Instance Segmentation
Box-supervised instance segmentation methods aim to achieve instance segmentation with only box annotations. Recent methods have demonstrated the effectiveness of acquiring high-quality pseudo masks under the teacher-stu…
Box-supervised Instance SegmentationInstance SegmentationSemantic SegmentationEnhancing Multimodal Large Language Models with Multi-instance Visual Prompt Generator for Visual Representation Enrichment
Multimodal Large Language Models (MLLMs) have achieved SOTA performance in various visual language tasks by fusing the visual representations with LLMs leveraging some visual adapters. In this paper, we first establish t…
Multimodal Information Interaction for Medical Image Segmentation
The use of multimodal data in assisted diagnosis and segmentation has emerged as a prominent area of interest in current research. However, one of the primary challenges is how to effectively fuse multimodal features. Mo…
Heart SegmentationImage SegmentationMedical Image SegmentationSegmentation+1