Enhancing Action Recognition by Leveraging the Hierarchical Structure of Actions and Textual Context
The sequential execution of actions and their hierarchical structure consisting of different levels of abstraction, provide features that remain unexplored in the task of action recognition. In this study, we present a novel approach to improve action recognition by exploiting the hierarchical organization of actions and by incorporating contextualized textual information, including location and prior actions to reflect the sequential context. To achieve this goal, we introduce a novel transformer architecture tailored for action recognition that utilizes both visual and textual features. Visual features are obtained from RGB and optical flow data, while text embeddings represent contextual information. Furthermore, we define a joint loss function to simultaneously train the model for both coarse and fine-grained action recognition, thereby exploiting the hierarchical nature of actions. To demonstrate the effectiveness of our method, we extend the Toyota Smarthome Untrimmed (TSU) dataset to introduce action hierarchies, introducing the Hierarchical TSU dataset. We also conduct an ablation study to assess the impact of different methods for integrating contextual and hierarchical data on action recognition performance. Results show that the proposed approach outperforms pre-trained SOTA methods when trained with the same hyperparameters. Moreover, they also show a 17.12% improvement in top-1 accuracy over the equivalent fine-grained RGB version when using ground-truth contextual information, and a 5.33% improvement when contextual information is obtained from actual predictions.
Code (1)
Tasks
Action RecognitionFine-grained Action RecognitionOptical Flow EstimationSimilar Papers 제목 키워드 기반
fine-CLIP: Enhancing Zero-Shot Fine-Grained Surgical Action Recognition with Vision-Language Models
While vision-language models like CLIP have advanced zero-shot surgical phase recognition, they struggle with fine-grained surgical activities, especially action triplets. This limitation arises because current CLIP form…
Action RecognitionSurgical phase recognitionTripletZero-Shot LearningMoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition
Audio-visual speech recognition (AVSR) has become critical for enhancing speech recognition in noisy environments by integrating both auditory and visual modalities. However, existing AVSR systems struggle to scale up wi…
Audio-Visual Speech RecognitionComputational EfficiencyMixture-of-ExpertsRobust Speech Recognition+3A Layer-Anchoring Strategy for Enhancing Cross-Lingual Speech Emotion Recognition
Cross-lingual speech emotion recognition (SER) is important for a wide range of everyday applications. While recent SER research relies heavily on large pretrained models for emotion training, existing studies often conc…
Emotion RecognitionSpeech Emotion RecognitionRFL: Simplifying Chemical Structure Recognition with Ring-Free Language
The primary objective of Optical Chemical Structure Recognition is to identify chemical structure images into corresponding markup sequences. However, the complex two-dimensional structures of molecules, particularly tho…
DecoderHierarchical Softmax for End-to-End Low-resource Multilingual Speech Recognition
Low-resource speech recognition has been long-suffering from insufficient training data. In this paper, we propose an approach that leverages neighboring languages to improve low-resource scenario performance, founded on…
speech-recognitionSpeech Recognition