DriverMHG: A Multi-Modal Dataset for Dynamic Recognition of Driver Micro Hand Gestures and a Real-Time Recognition Framework
The use of hand gestures provides a natural alternative to cumbersome interface devices for Human-Computer Interaction (HCI) systems. However, real-time recognition of dynamic micro hand gestures from video streams is challenging for in-vehicle scenarios since (i) the gestures should be performed naturally without distracting the driver, (ii) micro hand gestures occur within very short time intervals at spatially constrained areas, (iii) the performed gesture should be recognized only once, and (iv) the entire architecture should be designed lightweight as it will be deployed to an embedded system. In this work, we propose an HCI system for dynamic recognition of driver micro hand gestures, which can have a crucial impact in automotive sector especially for safety related issues. For this purpose, we initially collected a dataset named Driver Micro Hand Gestures (DriverMHG), which consists of RGB, depth and infrared modalities. The challenges for dynamic recognition of micro hand gestures have been addressed by proposing a lightweight convolutional neural network (CNN) based architecture which operates online efficiently with a sliding window approach. For the CNN model, several 3-dimensional resource efficient networks are applied and their performances are analyzed. Online recognition of gestures has been performed with 3D-MobileNetV2, which provided the best offline accuracy among the applied networks with similar computational complexities. The final architecture is deployed on a driver simulator operating in real-time. We make DriverMHG dataset and our source code publicly available.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
MultiMAE-DER: Multimodal Masked Autoencoder for Dynamic Emotion Recognition
This paper presents a novel approach to processing multimodal data for dynamic emotion recognition, named as the Multimodal Masked Autoencoder for Dynamic Emotion Recognition (MultiMAE-DER). The MultiMAE-DER leverages th…
Emotion RecognitionMultimodal Emotion RecognitionSelf-Supervised LearningVideo Emotion RecognitionImproving the Performance of Unimodal Dynamic Hand-Gesture Recognition with Multimodal Training
We present an efficient approach for leveraging the knowledge from multiple modalities in training unimodal 3D convolutional neural networks (3D-CNNs) for the task of dynamic hand gesture recognition. Instead of explicit…
Action RecognitionGesture RecognitionHand Gesture RecognitionHand-Gesture Recognition+1MM-DFN: Multimodal Dynamic Fusion Network for Emotion Recognition in Conversations
Emotion Recognition in Conversations (ERC) has considerable prospects for developing empathetic machines. For multimodal ERC, it is vital to understand context and fuse modality information in conversations. Recent graph…
Emotion RecognitionEmotion Recognition in ConversationA Deep Learning-based Multimodal Depth-Aware Dynamic Hand Gesture Recognition System
The dynamic hand gesture recognition task has seen studies on various unimodal and multimodal methods. Previously, researchers have explored depth and 2D-skeleton-based multimodal fusion CRNNs (Convolutional Recurrent Ne…
Gesture RecognitionHand Gesture RecognitionHand-Gesture RecognitionEMOE: Modality-Specific Enhanced Dynamic Emotion Experts
Multimodal Emotion Recognition (MER) aims to predict human emotions by leveraging multiple modalities, such as vision, acoustics, and language. However, due to the heterogeneity of these modalities, MER faces two key…
Emotion RecognitionIntent RecognitionMultimodal Emotion RecognitionMultimodal Intent Recognition