Motion Fused Frames: Data Level Fusion Strategy for Hand Gesture Recognition
Acquiring spatio-temporal states of an action is the most crucial step for action classification. In this paper, we propose a data level fusion strategy, Motion Fused Frames (MFFs), designed to fuse motion information into static images as better representatives of spatio-temporal states of an action. MFFs can be used as input to any deep learning architecture with very little modification on the network. We evaluate MFFs on hand gesture recognition tasks using three video datasets - Jester, ChaLearn LAP IsoGD and NVIDIA Dynamic Hand Gesture Datasets - which require capturing long-term temporal relations of hand movements. Our approach obtains very competitive performance on Jester and ChaLearn benchmarks with the classification accuracies of 96.28% and 57.4%, respectively, while achieving state-of-the-art performance with 84.7% accuracy on NVIDIA benchmark.
Code (1)
Tasks
Action ClassificationGeneral ClassificationGesture RecognitionHand Gesture RecognitionHand-Gesture RecognitionSimilar Papers 제목 키워드 기반
Frame Fusion with Vehicle Motion Prediction for 3D Object Detection
In LiDAR-based 3D detection, history point clouds contain rich temporal information helpful for future prediction. In the same way, history detections should contribute to future detections. In this paper, we propose a d…
3D Object DetectionFuture predictionModel Selectionmotion prediction+2Enhanced Fine-grained Motion Diffusion for Text-driven Human Motion Synthesis
The emergence of text-driven motion synthesis technique provides animators with great potential to create efficiently. However, in most cases, textual expressions only contain general and qualitative motion descriptions,…
Motion SynthesisvalidBodyFusion: Real-Time Capture of Human Motion and Surface Geometry Using a Single Depth Camera
We propose BodyFusion, a novel real-time geometry fusion method that can track and reconstruct non-rigid surface motion of a human performance using a single consumer-grade depth camera. To reduce the ambiguities of the …
Surface ReconstructionUncertainty-Weighted Image-Event Multimodal Fusion for Video Anomaly Detection
Most existing video anomaly detectors rely solely on RGB frames, which lack the temporal resolution needed to capture abrupt or transient motion cues, key indicators of anomalous events. To address this limitation, we pr…
Anomaly DetectionAnomaly Detection In Surveillance VideosVideo Anomaly DetectionVideo UnderstandingExplorative Inbetweening of Time and Space
We introduce bounded generation as a generalized task to control video generation to synthesize arbitrary camera and subject motion based only on a given start and end frame. Our objective is to fully leverage the inhere…
DenoisingVideo Generation