Do You Act Like You Talk? Exploring Pose-based Driver Action Classification with Speech Recognition Networks
Recognizing distractions on the road is crucial to reduce traffic accidents. Video-based networks are typically used, but are limited by their computational cost and are vulnerable to viewpoint changes. In this paper, we propose a novel approach for pose-based driver action classification using speech recognition networks, which is lighter and more viewpoint invariant that video-based one. We leverage the similarity in the encoding of information between audio and pose data, representing poses as key points over time. Our architecture is based on Squeezeformer, an efficient attentionbased speech recognition network. We introduce a selection of data augmentation techniques to enhance generalization. Experiments on the Drive&Act dataset demonstrate superior performance compared to state-of-the-art methods. Additionally, we explore the integration of object information and the impact of viewpoint changes. Our results highlight the effectiveness and robustness of speech recognition networks in pose-based action classification.
Code (1)
Tasks
Action ClassificationData AugmentationSkeleton Based Action Recognitionspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection
Active speaker detection (ASD) seeks to detect who is speaking in a visual scene of one or more speakers. The successful ASD depends on accurate interpretation of short-term and long-term audio and visual information, as…
Active Speaker DetectionAudio-Visual Active Speaker DetectionDrive-Net: Convolutional Network for Driver Distraction Detection
To help prevent motor vehicle accidents, there has been significant interest in finding an automated method to recognize signs of driver distraction, such as talking to passengers, fixing hair and makeup, eating and drin…
Multi-human Interactive Talking Dataset
Existing studies on talking video generation have predominantly focused on single-person monologues or isolated facial animations, limiting their applicability to realistic multi-human interactions. To bridge this gap, w…
Video GenerationPersonalized Autonomous Driving with Large Language Models: Field Experiments
Integrating large language models (LLMs) in autonomous vehicles enables conversation with AI systems to drive the vehicle. However, it also emphasizes the requirement for such systems to comprehend commands accurately an…
Autonomous DrivingAutonomous VehiclesLanguage ModellingLarge Language Model+2Our Cars Can Talk: How IoT Brings AI to Vehicles
Bringing AI to vehicles and enabling them as sensing platforms is key to transforming maintenance from reactive to proactive. Now is the time to integrate AI copilots that speak both languages: machine and driver. This a…