paper-with-me

홈 › Papers

Enhancing Action Recognition by Leveraging the Hierarchical Structure of Actions and Textual Context

2024-10-28 · Manuel Benavent-Lledo, David Mulero-Pérez, David Ortiz-Perez, Jose Garcia-Rodriguez, Antonis Argyros

The sequential execution of actions and their hierarchical structure consisting of different levels of abstraction, provide features that remain unexplored in the task of action recognition. In this study, we present a novel approach to improve action recognition by exploiting the hierarchical organization of actions and by incorporating contextualized textual information, including location and prior actions to reflect the sequential context. To achieve this goal, we introduce a novel transformer architecture tailored for action recognition that utilizes both visual and textual features. Visual features are obtained from RGB and optical flow data, while text embeddings represent contextual information. Furthermore, we define a joint loss function to simultaneously train the model for both coarse and fine-grained action recognition, thereby exploiting the hierarchical nature of actions. To demonstrate the effectiveness of our method, we extend the Toyota Smarthome Untrimmed (TSU) dataset to introduce action hierarchies, introducing the Hierarchical TSU dataset. We also conduct an ablation study to assess the impact of different methods for integrating contextual and hierarchical data on action recognition performance. Results show that the proposed approach outperforms pre-trained SOTA methods when trained with the same hyperparameters. Moreover, they also show a 17.12% improvement in top-1 accuracy over the equivalent fine-grained RGB version when using ground-truth contextual information, and a 5.33% improvement when contextual information is obtained from actual predictions.

📄 PDF Abstract BibTeX arXiv:2410.21275

Code (1)

3dperceptionlab/hierarchicalactionrecognition 공식 구현 pytorch

Tasks

Action RecognitionFine-grained Action RecognitionOptical Flow Estimation

Similar Papers 제목 키워드 기반

fine-CLIP: Enhancing Zero-Shot Fine-Grained Surgical Action Recognition with Vision-Language Models

2025-03-25 · Saurav Sharma, Didier Mutter, Nicolas Padoy

While vision-language models like CLIP have advanced zero-shot surgical phase recognition, they struggle with fine-grained surgical activities, especially action triplets. This limitation arises because current CLIP form…

Action RecognitionSurgical phase recognitionTripletZero-Shot Learning

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition

2025-02-11 · Sungnyun Kim, Kangwook Jang, Sangmin Bae, Sungwoo Cho 외

Audio-visual speech recognition (AVSR) has become critical for enhancing speech recognition in noisy environments by integrating both auditory and visual modalities. However, existing AVSR systems struggle to scale up wi…

Audio-Visual Speech RecognitionComputational EfficiencyMixture-of-ExpertsRobust Speech Recognition+3

A Layer-Anchoring Strategy for Enhancing Cross-Lingual Speech Emotion Recognition

2024-07-06 · Shreya G. Upadhyay, Carlos Busso, Chi-Chun Lee

Cross-lingual speech emotion recognition (SER) is important for a wide range of everyday applications. While recent SER research relies heavily on large pretrained models for emotion training, existing studies often conc…

Emotion RecognitionSpeech Emotion Recognition

RFL: Simplifying Chemical Structure Recognition with Ring-Free Language

2024-12-10 · Qikai Chang, Mingjun Chen, Changpeng Pi, Pengfei Hu 외

The primary objective of Optical Chemical Structure Recognition is to identify chemical structure images into corresponding markup sequences. However, the complex two-dimensional structures of molecules, particularly tho…

Decoder

Hierarchical Softmax for End-to-End Low-resource Multilingual Speech Recognition

2022-04-08 · Qianying Liu, Zhuo Gong, Zhengdong Yang, Yuhang Yang 외

Low-resource speech recognition has been long-suffering from insufficient training data. In this paper, we propose an approach that leverages neighboring languages to improve low-resource scenario performance, founded on…

speech-recognitionSpeech Recognition