Zero-Shot Action Recognition
7개 벤치마크 · 논문 93편 · 이 태스크의 논문 보기 →
Benchmarks
UCF101
HMDB51
Kinetics
Olympics
ActivityNet
Charades
THUMOS' 14
Most implemented
LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
Bidirectional Cross-Modal Knowledge Exploration for Video Recognition with Pre-trained Vision-Language Models
EVA-CLIP: Improved Training Techniques for CLIP at Scale
Learning a Deep Embedding Model for Zero-Shot Learning
Leveraging Temporal Contextualization for Video Action Recognition
Papers
Divide, Deliberate, Decide: A Multi-Agent Framework for Fine-Grained Egocentric Action Recognition
Fine-grained action recognition in egocentric video is challenging for Vision-Language Models (VLMs): actions often differ only in small visual cues, and a single model tends to be biased toward a subset of these cues. W…
Zero-Shot Action RecognitionCross-Domain Human Action Recognition from Multiview Motion and Textual Descriptions
Robustness to domain changes is a key capability for effective deployment of human action recognition systems in real-world scenarios, where action categories at inference can present important domain shifts or even unse…
Zero-Shot Action RecognitionTransfer LearningCEZSAR: A Contrastive Embedding Method for Zero-Shot Action Recognition
This paper proposes a novel Zero-Shot Action Recognition~(ZSAR) method based on contrastive learning. In ZSAR, we aim to classify examples from classes that were missing during training. Two well-known problems remain in…
Zero-Shot Action RecognitionContrastive LearningMotion-Guided Semantic Alignment with Negative Prompts for Zero-Shot Video Action Recognition
Zero-shot action recognition is challenging due to the semantic gap between seen and unseen classes. We present a novel framework that enhances CLIP with disentangled embeddings and semantic-guided interaction. A Motion …
Zero-Shot Action RecognitionNovel Semantic Prompting for Zero-Shot Action Recognition
Zero-shot action recognition relies on transferring knowledge from vision-language models to unseen actions using semantic descriptions. While recent methods focus on temporal modeling or architectural adaptations to han…
Zero-Shot Action RecognitionAction UnderstandingDynaPURLS: Dynamic Refinement of Part-Aware Representations for Skeleton-Based Zero-Shot Action Recognition
Zero-shot skeleton-based action recognition (ZS-SAR) is fundamentally constrained by prevailing approaches that rely on aligning skeleton features with static, class-level semantics. This coarse-grained alignment fails t…
Zero-Shot Action Recognition