paper-with-me

홈 › Papers

Multi-modal Fusion for Single-Stage Continuous Gesture Recognition

2020-11-10 · Harshala Gammulle, Simon Denman, Sridha Sridharan, Clinton Fookes

Gesture recognition is a much studied research area which has myriad real-world applications including robotics and human-machine interaction. Current gesture recognition methods have focused on recognising isolated gestures, and existing continuous gesture recognition methods are limited to two-stage approaches where independent models are required for detection and classification, with the performance of the latter being constrained by detection performance. In contrast, we introduce a single-stage continuous gesture recognition framework, called Temporal Multi-Modal Fusion (TMMF), that can detect and classify multiple gestures in a video via a single model. This approach learns the natural transitions between gestures and non-gestures without the need for a pre-processing segmentation step to detect individual gestures. To achieve this, we introduce a multi-modal fusion mechanism to support the integration of important information that flows from multi-modal inputs, and is scalable to any number of modes. Additionally, we propose Unimodal Feature Mapping (UFM) and Multi-modal Feature Mapping (MFM) models to map uni-modal features and the fused multi-modal features respectively. To further enhance performance, we propose a mid-point based loss function that encourages smooth alignment between the ground truth and the prediction, helping the model to learn natural gesture transitions. We demonstrate the utility of our proposed framework, which can handle variable-length input videos, and outperforms the state-of-the-art on three challenging datasets: EgoGesture, IPN hand, and ChaLearn LAP Continuous Gesture Dataset (ConGD). Furthermore, ablation experiments show the importance of different components of the proposed framework.

📄 PDF Abstract BibTeX arXiv:2011.04945

Code (0)

등록된 구현이 없습니다.

Tasks

Gesture Recognition

Similar Papers 제목 키워드 기반

Stage-Adaptive Reliability Modeling for Continuous Valence-Arousal Estimation

2026-03-12 · Yubeen Lee, Sangeun Lee, Junyeop Cha, Eunil Park arxiv

Continuous valence-arousal estimation in real-world environments is challenging due to inconsistent modality reliability and interaction-dependent variability in audio-visual signals. Existing approaches primarily focus …

AVT2-DWF: Improving Deepfake Detection with Audio-Visual Fusion and Dynamic Weighting Strategies

2024-03-22 · Rui Wang, Dengpan Ye, Long Tang, Yunming Zhang 외

With the continuous improvements of deepfake methods, forgery messages have transitioned from single-modality to multi-modal fusion, posing new challenges for existing forgery detection algorithms. In this paper, we prop…

DeepFake DetectionFace Swapping

MMDR: A Result Feature Fusion Object Detection Approach for Autonomous System

2023-04-19 · Wendong Zhang

Object detection has been extensively utilized in autonomous systems in recent years, encompassing both 2D and 3D object detection. Recent research in this field has primarily centered around multimodal approaches for ad…

3D Object DetectionObjectobject-detectionObject Detection

DAgger Diffusion Navigation: DAgger Boosted Diffusion Policy for Vision-Language Navigation

2025-08-13 · Haoxiang Shi, Xiang Deng, Zaijing Li, Gongwei Chen 외 arxiv

Vision-Language Navigation in Continuous Environments (VLN-CE) requires agents to follow natural language instructions through free-form 3D spaces. Existing VLN-CE approaches typically use a two-stage waypoint planning f…

Vision-Language NavigationSpatial Reasoning

Pay "Attention" to Adverse Weather: Weather-aware Attention-based Object Detection

2022-04-22 · Saket S. Chaturvedi, Lan Zhang, Xiaoyong Yuan

Despite the recent advances of deep neural networks, object detection for adverse weather remains challenging due to the poor perception of some sensors in adverse weather. Instead of relying on one single sensor, multim…

object-detectionObject Detection