paper-with-me

Papers

MultiFuser: Multimodal Fusion Transformer for Enhanced Driver Action Recognition

2024-08-03 · Ruoyu Wang, Wenqian Wang, Jianjun Gao, Dan Lin, Kim-Hui Yap, Bingbing Li

Driver action recognition, aiming to accurately identify drivers' behaviours, is crucial for enhancing driver-vehicle interactions and ensuring driving safety. Unlike general action recognition, drivers' environments are often challenging, being gloomy and dark, and with the development of sensors, various cameras such as IR and depth cameras have emerged for analyzing drivers' behaviors. Therefore, in this paper, we propose a novel multimodal fusion transformer, named MultiFuser, which identifies cross-modal interrelations and interactions among multimodal car cabin videos and adaptively integrates different modalities for improved representations. Specifically, MultiFuser comprises layers of Bi-decomposed Modules to model spatiotemporal features, with a modality synthesizer for multimodal features integration. Each Bi-decomposed Module includes a Modal Expertise ViT block for extracting modality-specific features and a Patch-wise Adaptive Fusion block for efficient cross-modal fusion. Extensive experiments are conducted on Drive&Act dataset and the results demonstrate the efficacy of our proposed approach.

📄 PDF Abstract BibTeX arXiv:2408.01766

Code (0)

등록된 구현이 없습니다.

Tasks

Action Recognition

Similar Papers 제목 키워드 기반

Robust Multiview Multimodal Driver Monitoring System Using Masked Multi-Head Self-Attention

2023-04-13 · Yiming Ma, Victor Sanchez, Soodeh Nikan, Devesh Upadhyay 외

Driver Monitoring Systems (DMSs) are crucial for safe hand-over actions in Level-2+ self-driving vehicles. State-of-the-art DMSs leverage multiple sensors mounted at different locations to monitor the driver and the vehi…

Contrastive LearningGPU

DiffAttn: Diffusion-Based Drivers' Visual Attention Prediction with LLM-Enhanced Semantic Reasoning

2026-03-30 · Weimin Liu, Qingkun Li, Jiyuan Qiu, Wenjun Wang 외 arxiv

Drivers' visual attention provides critical cues for anticipating latent hazards and directly shapes decision-making and control maneuvers, where its absence can compromise traffic safety. To emulate drivers' perception …

Scene Understanding

RoboFlamingo-Plus: Fusion of Depth and RGB Perception with Vision-Language Models for Enhanced Robotic Manipulation

2025-03-25 · Sheng Wang

As robotic technologies advancing towards more complex multimodal interactions and manipulation tasks, the integration of advanced Vision-Language Models (VLMs) has become a key driver in the field. Despite progress with…

Audio-Visual Speech Recognition based on Regulated Transformer and Spatio-Temporal Fusion Strategy for Driver Assistive Systems

2024-05-09 · Expert Systems with Applications 2024 5 · Dmitry Ryumin, Alexandr Axyonov, Elena Ryumina, Denis Ivanko 외

This article presents a research methodology for audio-visual speech recognition (AVSR) in driver assistive systems. These systems necessitate ongoing interaction with drivers while driving through voice control for safe…

Audio-Visual Speech RecognitionLipreadingLip Readingspeech-recognition+2

A Comparative Analysis of Decision-Level Fusion for Multimodal Driver Behaviour Understanding

2022-04-10 · Alina Roitberg, Kunyu Peng, Zdravko Marinov, Constantin Seibold 외

Visual recognition inside the vehicle cabin leads to safer driving and more intuitive human-vehicle interaction but such systems face substantial obstacles as they need to capture different granularities of driver behavi…