paper-with-me

Papers

Using Multimodal Large Language Models for Automated Detection of Traffic Safety Critical Events

2024-06-19 · Mohammad Abu Tami, Huthaifa I. Ashqar, Mohammed Elhenawy

Traditional approaches to safety event analysis in autonomous systems have relied on complex machine learning models and extensive datasets for high accuracy and reliability. However, the advent of Multimodal Large Language Models (MLLMs) offers a novel approach by integrating textual, visual, and audio modalities, thereby providing automated analyses of driving videos. Our framework leverages the reasoning power of MLLMs, directing their output through context-specific prompts to ensure accurate, reliable, and actionable insights for hazard detection. By incorporating models like Gemini-Pro-Vision 1.5 and Llava, our methodology aims to automate the safety critical events and mitigate common issues such as hallucinations in MLLM outputs. Preliminary results demonstrate the framework's potential in zero-shot learning and accurate scenario analysis, though further validation on larger datasets is necessary. Furthermore, more investigations are required to explore the performance enhancements of the proposed framework through few-shot learning and fine-tuned models. This research underscores the significance of MLLMs in advancing the analysis of the naturalistic driving videos by improving safety-critical event detecting and understanding the interaction with complex environments.

📄 PDF Abstract BibTeX arXiv:2406.13894

Code (0)

등록된 구현이 없습니다.

Tasks

Few-Shot LearningZero-Shot Learning

Similar Papers 제목 키워드 기반

Investigating Traffic Accident Detection Using Multimodal Large Language Models

2025-09-23 · Ilhan Skender, Kailin Tong, Selim Solmaz, Daniel Watzenig arxiv

Traffic safety remains a critical global concern, with timely and accurate accident detection essential for hazard reduction and rapid emergency response. Infrastructure-based vision sensors offer scalable and efficient …

Traffic Accident DetectionMulti-Object TrackingInstance SegmentationObject Detection

AITP: Traffic Accident Responsibility Allocation via Multimodal Large Language Models

2026-04-11 · Zijin Zhou, Songan Zhang arxiv

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in Traffic Accident Detection (TAD) and Traffic Accident Understanding (TAU). However, existing studies mainly focus on describing and interpreti…

Traffic Accident Detection

TrafficRAG: A Multimodal RAG Framework for Traffic Accident Liability Determination

2026-06-01 · Xu Li, Zedong Fu, Xinyi Li, Xun Han arxiv

Traffic accident liability analysis is a critical yet challenging task in intelligent transportation and legal assistance. Existing methods often suffer from low efficiency, subjective judgment, and inconsistent analysis…

When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis

2025-01-17 · Ruixuan Zhang, Beichen Wang, Juexiao Zhang, Zilin Bian 외

The increasing availability of traffic videos functioning on a 24/7/365 time scale has the great potential of increasing the spatio-temporal coverage of traffic accidents, which will help improve traffic safety. However,…

Large Language ModelMultimodal Large Language Modelobject-detectionObject Detection+2

Detection of Illicit Drug Trafficking Events on Instagram: A Deep Multimodal Multilabel Learning Approach

2021-08-19 · Chuanbo Hu, Minglei Yin, Bin Liu, Xin Li 외

Social media such as Instagram and Twitter have become important platforms for marketing and selling illicit drugs. Detection of online illicit drug trafficking has become critical to combat the online trade of illicit d…

Marketing