paper-with-me

Papers

Enhancing Multimodal Large Language Models for Safety-Critical Driving Video Analysis

2026-05-21 · Tomaso Trinci, Henrique Piñeiro Monteagudo, Leonardo Taccari arxiv

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in general visual understanding. However, their application to safety-critical driving scenarios remains limited by an inability to accurately perceive and reason about rare high-stakes dynamic events, such as collisions or near-collisions. To address this, we introduce a pipeline that enhances MLLM perception by fusing downsampled video frames with synchronized high-frequency telematics data (IMU and GPS) and semantic insights from specialized computer vision models. Our pipeline generates high-quality pseudo-labels, including descriptive captions and question-answer pairs, specifically designed to train MLLMs to identify and describe Safety-Critical Events (SCEs) in real-world driving footage. We show the effectiveness of our approach fine-tuning the open-source QwenVL-2.5 model via DoRA adapters: our experiments demonstrate significant improvements in identifying and explaining safety-critical events, with fewer than 50M trainable parameters and limited computational budget.

📄 PDF Abstract BibTeX arXiv:2605.22185

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models

2025-06-02 · Youze Wang, WenBo Hu, Yinpeng Dong, Jing Liu 외

Large Language Models (LLMs) have evolved into Multimodal Large Language Models (MLLMs), significantly enhancing their capabilities by integrating visual information and other types, thus aligning more closely with the n…

Safety Alignment

Multimodal Large Language Models for Enhanced Traffic Safety: A Comprehensive Review and Future Trends

2025-04-21 · Mohammad Abu Tami, Mohammed Elhenawy, Huthaifa I. Ashqar

Traffic safety remains a critical global challenge, with traditional Advanced Driver-Assistance Systems (ADAS) often struggling in dynamic real-world scenarios due to fragmented sensor processing and susceptibility to ad…

Adversarial RobustnessDecision MakingScene Understanding

Automated Hazard Detection in Construction Sites Using Large Language and Vision-Language Models

2025-11-13 · Islem Sahraoui arxiv

This thesis explores a multimodal AI framework for enhancing construction safety through the combined analysis of textual and visual data. In safety-critical environments such as construction sites, accident data often e…

OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models

2025-11-13 · Yuping Yan, Yuhan Xie, Yuanshuai Li, Yingchao Yu 외 arxiv

Since Multimodal Large Language Models (MLLMs) are increasingly being integrated into everyday tools and intelligent agents, growing concerns have arisen regarding their possible output of unsafe contents, ranging from t…

Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing

2024-09-26 · Huthaifa I. Ashqar, Ahmed Jaber, Taqwa I. Alhadidi, Mohammed Elhenawy

This study aims to comprehensively review and empirically evaluate the application of multimodal large language models (MLLMs) and Large Vision Models (VLMs) in object detection for transportation systems. In the first f…

Event DetectionObjectobject-detectionObject Detection+1