paper-with-me

홈 › Papers

Unifying Few- and Zero-Shot Egocentric Action Recognition

2020-05-27 · Scott Tyler R., Shvartsman Michael, Ridgeway Karl

Although there has been significant research in egocentric action recognition, most methods and tasks, including EPIC-KITCHENS, suppose a fixed set of action classes. Fixed-set classification is useful for benchmarking methods, but is often unrealistic in practical settings due to the compositionality of actions, resulting in a functionally infinite-cardinality label set. In this work, we explore generalization with an open set of classes by unifying two popular approaches: few- and zero-shot generalization (the latter which we reframe as cross-modal few-shot generalization). We propose a new set of splits derived from the EPIC-KITCHENS dataset that allow evaluation of open-set classification, and use these splits to show that adding a metric-learning loss to the conventional direct-alignment baseline can improve zero-shot classification by as much as 10%, while not sacrificing few-shot performance.

📄 PDF Abstract BibTeX arXiv:2006.11393

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionBenchmarkingClassificationGeneral ClassificationMetric Learningopen-set classificationzero-shot-classificationZero-shot GeneralizationZero-Shot Learning

Similar Papers 제목 키워드 기반

GPT4Ego: Unleashing the Potential of Pre-trained Models for Zero-Shot Egocentric Action Recognition

2024-01-18 · Guangzhao Dai, Xiangbo Shu, Wenhao Wu, Rui Yan 외

Vision-Language Models (VLMs), pre-trained on large-scale datasets, have shown impressive performance in various visual recognition tasks. This advancement paves the way for notable performance in Zero-Shot Egocentric Ac…

Action RecognitionText Matching

Divide, Deliberate, Decide: A Multi-Agent Framework for Fine-Grained Egocentric Action Recognition

2026-06-16 · Alessandro Sottovia, Alessandro Torcinovich, Oswald Lanz arxiv

Fine-grained action recognition in egocentric video is challenging for Vision-Language Models (VLMs): actions often differ only in small visual cues, and a single model tends to be biased toward a subset of these cues. W…

Zero-Shot Action Recognition

LLM as A Robotic Brain: Unifying Egocentric Memory and Control

2023-04-19 · Jinjie Mai, Jun Chen, Bing Li, Guocheng Qian 외

Embodied AI focuses on the study and development of intelligent systems that possess a physical or virtual embodiment (i.e. robots) and are able to dynamically interact with their environment. Memory and control are the …

Embodied Question AnsweringLanguage ModelingLanguage ModellingQuestion Answering+1

X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization

2024-03-28 · CVPR 2024 1 · Anna Kukleva, Fadime Sener, Edoardo Remelli, Bugra Tekin 외

Lately, there has been growing interest in adapting vision-language models (VLMs) to image and third-person video classification due to their success in zero-shot recognition. However, the adaptation of these models to e…

Video ClassificationZero-Shot Learning

DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation

2025-09-14 · Yunheng Wang, Yuetong Fang, Taowen Wang, Yixiao Feng 외 arxiv

Vision-and-Language Navigation in Continuous Environments (VLN-CE), which links language instructions to perception and control in the real world, is a core capability of embodied robots. Recently, large-scale pretrained…

Scene Understanding