paper-with-me

홈 › Papers

When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis

2025-01-17 · Ruixuan Zhang, Beichen Wang, Juexiao Zhang, Zilin Bian, Chen Feng, Kaan Ozbay

The increasing availability of traffic videos functioning on a 24/7/365 time scale has the great potential of increasing the spatio-temporal coverage of traffic accidents, which will help improve traffic safety. However, analyzing footage from hundreds, if not thousands, of traffic cameras in a 24/7/365 working protocol remains an extremely challenging task, as current vision-based approaches primarily focus on extracting raw information, such as vehicle trajectories or individual object detection, but require laborious post-processing to derive actionable insights. We propose SeeUnsafe, a new framework that integrates Multimodal Large Language Model (MLLM) agents to transform video-based traffic accident analysis from a traditional extraction-then-explanation workflow to a more interactive, conversational approach. This shift significantly enhances processing throughput by automating complex tasks like video classification and visual grounding, while improving adaptability by enabling seamless adjustments to diverse traffic scenarios and user-defined queries. Our framework employs a severity-based aggregation strategy to handle videos of various lengths and a novel multimodal prompt to generate structured responses for review and evaluation and enable fine-grained visual grounding. We introduce IMS (Information Matching Score), a new MLLM-based metric for aligning structured responses with ground truth. We conduct extensive experiments on the Toyota Woven Traffic Safety dataset, demonstrating that SeeUnsafe effectively performs accident-aware video classification and visual grounding by leveraging off-the-shelf MLLMs. Source code will be available at \url{https://github.com/ai4ce/SeeUnsafe}.

📄 PDF Abstract BibTeX arXiv:2501.10604

Code (1)

ai4ce/seeunsafe 공식 구현 pytorch

Tasks

Large Language ModelMultimodal Large Language Modelobject-detectionObject DetectionVideo ClassificationVisual Grounding

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

EG-ARSA: An Expert-Grounded Open Model for Visual Road Safety Auditing in Low-Resource Settings

2026-08-24 · Md Thamed Bin Zaman Chowdhury, Moazzem Hossain arxiv

Road traffic injuries remain a major challenge in low- and middle-income countries, where proactive road safety auditing is limited by incomplete crash records, shortages of qualified auditors, and the high cost of large…

Understanding Real-World Traffic Safety through RoadSafe365 Benchmark

2026-02-06 · Xinyu Liu, Darryl C. Jacob, Yuxin Liu, Xinsong Du 외 arxiv

Although recent traffic benchmarks have advanced multimodal data analysis, they generally lack systematic evaluation aligned with official safety standards. To fill this gap, we introduce RoadSafe365, a large-scale visio…

Watchdogs and Oracles: Runtime Verification Meets Large Language Models for Autonomous Systems

2025-11-18 · Angelo Ferrando arxiv

Assuring the safety and trustworthiness of autonomous systems is particularly difficult when learning-enabled components and open environments are involved. Formal methods provide strong guarantees but depend on complete…

Vision-Language Models for Autonomous Driving: CLIP-Based Dynamic Scene Understanding

2025-01-09 · Mohammed Elhenawy, Huthaifa I. Ashqar, Andry Rakotonirainy, Taqwa I. Alhadidi 외

Scene understanding is essential for enhancing driver safety, generating human-centric explanations for Automated Vehicle (AV) decisions, and leveraging Artificial Intelligence (AI) for retrospective driving video analys…

Autonomous DrivingIn-Context LearningScene ClassificationScene Recognition+1

Are Vision LLMs Road-Ready? A Comprehensive Benchmark for Safety-Critical Driving Video Understanding

2025-04-20 · Tong Zeng, Longfeng Wu, Liang Shi, Dawei Zhou 외

Vision Large Language Models (VLLMs) have demonstrated impressive capabilities in general visual tasks such as image captioning and visual question answering. However, their effectiveness in specialized, safety-critical …

Autonomous DrivingImage CaptioningMultiple-choiceQuestion Answering+3