Developing Motion Code Embedding for Action Recognition in Videos
In this work, we propose a motion embedding strategy known as motion codes, which is a vectorized representation of motions based on a manipulation's salient mechanical attributes. These motion codes provide a robust motion representation, and they are obtained using a hierarchy of features called the motion taxonomy. We developed and trained a deep neural network model that combines visual and semantic features to identify the features found in our motion taxonomy to embed or annotate videos with motion codes. To demonstrate the potential of motion codes as features for machine learning tasks, we integrated the extracted features from the motion embedding model into the current state-of-the-art action recognition model. The obtained model achieved higher accuracy than the baseline model for the verb classification task on egocentric videos from the EPIC-KITCHENS dataset.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionAction Recognition In VideosSimilar Papers 제목 키워드 기반
Take an Emotion Walk: Perceiving Emotions from Gaits Using Hierarchical Attention Pooling and Affective Mapping
We present an autoencoder-based semi-supervised approach to classify perceived human emotions from walking styles obtained from videos or motion-captured data and represented as sequences of 3D poses. Given the motion on…
Action RecognitionDecoderEmotion RecognitionHICEM: A High-Coverage Emotion Model for Artificial Emotional Intelligence
As social robots and other intelligent machines enter the home, artificial emotional intelligence (AEI) is taking center stage to address users' desire for deeper, more meaningful human-machine interaction. To accomplish…
DescriptiveEmotional IntelligenceEmotion RecognitionVocal Bursts Intensity Prediction+1Continuous Metric Learning For Transferable Speech Emotion Recognition and Embedding Across Low-resource Languages
Speech emotion recognition~(SER) refers to the technique of inferring the emotional state of an individual from speech signals. SERs continue to garner interest due to their wide applicability. Although the domain is mai…
DenoisingEmotion ClassificationEmotion RecognitionMetric Learning+1Motion-Guided Semantic Alignment with Negative Prompts for Zero-Shot Video Action Recognition
Zero-shot action recognition is challenging due to the semantic gap between seen and unseen classes. We present a novel framework that enhances CLIP with disentangled embeddings and semantic-guided interaction. A Motion …
Zero-Shot Action RecognitionSL-DML: Signal Level Deep Metric Learning for Multimodal One-Shot Action Recognition
Recognizing an activity with a single reference sample using metric learning approaches is a promising research field. The majority of few-shot methods focus on object recognition or face-identification. We propose a met…
Action RecognitionFace IdentificationMetric LearningObject Recognition+2