Learning Using Privileged Information for Zero-Shot Action Recognition
Zero-Shot Action Recognition (ZSAR) aims to recognize video actions that have never been seen during training. Most existing methods assume a shared semantic space between seen and unseen actions and intend to directly learn a mapping from a visual space to the semantic space. This approach has been challenged by the semantic gap between the visual space and semantic space. This paper presents a novel method that uses object semantics as privileged information to narrow the semantic gap and, hence, effectively, assist the learning. In particular, a simple hallucination network is proposed to implicitly extract object semantics during testing without explicitly extracting objects and a cross-attention module is developed to augment visual feature with the object semantics. Experiments on the Olympic Sports, HMDB51 and UCF101 datasets have shown that the proposed method outperforms the state-of-the-art methods by a large margin.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionHallucinationObjectZero-Shot Action RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Privileged Zero-Shot AutoML
This work improves the quality of automated machine learning (AutoML) systems by using dataset and function descriptions while significantly decreasing computation time from minutes to milliseconds by using a zero-shot a…
AutoMLBIG-bench Machine LearningGraph Neural NetworkRepresentation LearningAlternative Semantic Representations for Zero-Shot Human Action Recognition
A proper semantic representation for encoding side information is key to the success of zero-shot learning. In this paper, we explore two alternative semantic representations especially for zero-shot human action recogni…
Action RecognitionTemporal Action LocalizationZero-Shot Action RecognitionZero-Shot LearningA New Split for Evaluating True Zero-Shot Action Recognition
Zero-shot action recognition is the task of classifying action categories that are not available in the training set. In this setting, the standard evaluation protocol is to use existing action recognition datasets(e.g. …
Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionZero-Shot Action Recognition+1Symbiotic Attention with Privileged Information for Egocentric Action Recognition
Egocentric video recognition is a natural testbed for diverse interaction reasoning. Due to the large action vocabulary in egocentric video datasets, recent studies usually utilize a two-branch structure for action recog…
Action RecognitionEgocentric Activity RecognitionGeneral Classificationobject-detection+3Learning with Privileged Information for Multi-Label Classification
In this paper, we propose a novel approach for learning multi-label classifiers with the help of privileged information. Specifically, we use similarity constraints to capture the relationship between available informati…
Action Unit DetectionClassificationFacial Action Unit DetectionGeneral Classification+4