paper-with-me

Papers

Leveraging Multimodal-LLMs Assisted by Instance Segmentation for Intelligent Traffic Monitoring

2025-02-16 · Murat Arda Onsu, Poonam Lohan, Burak Kantarci, Aisha Syed, Matthew Andrews, Sean Kennedy

A robust and efficient traffic monitoring system is essential for smart cities and Intelligent Transportation Systems (ITS), using sensors and cameras to track vehicle movements, optimize traffic flow, reduce congestion, enhance road safety, and enable real-time adaptive traffic control. Traffic monitoring models must comprehensively understand dynamic urban conditions and provide an intuitive user interface for effective management. This research leverages the LLaVA visual grounding multimodal large language model (LLM) for traffic monitoring tasks on the real-time Quanser Interactive Lab simulation platform, covering scenarios like intersections, congestion, and collisions. Cameras placed at multiple urban locations collect real-time images from the simulation, which are fed into the LLaVA model with queries for analysis. An instance segmentation model integrated into the cameras highlights key elements such as vehicles and pedestrians, enhancing training and throughput. The system achieves 84.3% accuracy in recognizing vehicle locations and 76.4% in determining steering direction, outperforming traditional models.

📄 PDF Abstract BibTeX arXiv:2502.11304

Code (0)

등록된 구현이 없습니다.

Tasks

Instance SegmentationLanguage ModelingLanguage ModellingLarge Language ModelMultimodal Large Language ModelSemantic SegmentationVisual Grounding

Similar Papers 제목 키워드 기반

MLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge

2026-08-24 · Liangtao Shi, Jinxia Xie, Xiantao Hu, Ting Liu arxiv

In this technical report, we present a training-free framework for audio-guided video object segmentation, which integrates Multimodal Large Language Models (MLLMs) with SAM-based segmentation models. We decompose the ta…

Video Object SegmentationMultimodal ReasoningVideo Segmentation

All in One: Visual-Description-Guided Unified Point Cloud Segmentation

2025-07-07 · Zongyan Han, Mohamed El Amine Boudjoghra, Jiahua Dong, Jinhong Wang 외 arxiv

Unified segmentation of 3D point clouds is crucial for scene understanding, but is hindered by its sparse structure, limited annotations, and the challenge of distinguishing fine-grained object classes in complex environ…

Point Cloud SegmentationPanoptic SegmentationScene UnderstandingPoint Clouds

BoxSeg: Quality-Aware and Peer-Assisted Learning for Box-supervised Instance Segmentation

2025-04-07 · Jinxiang Lai, Wenlong Wu, Jiawei Zhan, Jian Li 외

Box-supervised instance segmentation methods aim to achieve instance segmentation with only box annotations. Recent methods have demonstrated the effectiveness of acquiring high-quality pseudo masks under the teacher-stu…

Box-supervised Instance SegmentationInstance SegmentationSemantic Segmentation

Enhancing Multimodal Large Language Models with Multi-instance Visual Prompt Generator for Visual Representation Enrichment

2024-06-05 · Wenliang Zhong, Wenyi Wu, Qi Li, Rob Barton 외

Multimodal Large Language Models (MLLMs) have achieved SOTA performance in various visual language tasks by fusing the visual representations with LLMs leveraging some visual adapters. In this paper, we first establish t…

Multimodal Information Interaction for Medical Image Segmentation

2024-04-25 · Xinxin Fan, Lin Liu, Haoran Zhang

The use of multimodal data in assisted diagnosis and segmentation has emerged as a prominent area of interest in current research. However, one of the primary challenges is how to effectively fuse multimodal features. Mo…

Heart SegmentationImage SegmentationMedical Image SegmentationSegmentation+1