Unifying Few- and Zero-Shot Egocentric Action Recognition
Although there has been significant research in egocentric action recognition, most methods and tasks, including EPIC-KITCHENS, suppose a fixed set of action classes. Fixed-set classification is useful for benchmarking methods, but is often unrealistic in practical settings due to the compositionality of actions, resulting in a functionally infinite-cardinality label set. In this work, we explore generalization with an open set of classes by unifying two popular approaches: few- and zero-shot generalization (the latter which we reframe as cross-modal few-shot generalization). We propose a new set of splits derived from the EPIC-KITCHENS dataset that allow evaluation of open-set classification, and use these splits to show that adding a metric-learning loss to the conventional direct-alignment baseline can improve zero-shot classification by as much as 10%, while not sacrificing few-shot performance.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionBenchmarkingClassificationGeneral ClassificationMetric Learningopen-set classificationzero-shot-classificationZero-shot GeneralizationZero-Shot LearningSimilar Papers 제목 키워드 기반
GPT4Ego: Unleashing the Potential of Pre-trained Models for Zero-Shot Egocentric Action Recognition
Vision-Language Models (VLMs), pre-trained on large-scale datasets, have shown impressive performance in various visual recognition tasks. This advancement paves the way for notable performance in Zero-Shot Egocentric Ac…
Action RecognitionText MatchingDivide, Deliberate, Decide: A Multi-Agent Framework for Fine-Grained Egocentric Action Recognition
Fine-grained action recognition in egocentric video is challenging for Vision-Language Models (VLMs): actions often differ only in small visual cues, and a single model tends to be biased toward a subset of these cues. W…
Zero-Shot Action RecognitionLLM as A Robotic Brain: Unifying Egocentric Memory and Control
Embodied AI focuses on the study and development of intelligent systems that possess a physical or virtual embodiment (i.e. robots) and are able to dynamically interact with their environment. Memory and control are the …
Embodied Question AnsweringLanguage ModelingLanguage ModellingQuestion Answering+1X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization
Lately, there has been growing interest in adapting vision-language models (VLMs) to image and third-person video classification due to their success in zero-shot recognition. However, the adaptation of these models to e…
Video ClassificationZero-Shot LearningDreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation
Vision-and-Language Navigation in Continuous Environments (VLN-CE), which links language instructions to perception and control in the real world, is a core capability of embodied robots. Recently, large-scale pretrained…
Scene Understanding