Open Vocabulary Action Recognition
2개 벤치마크 · 논문 6편 · 이 태스크의 논문 보기 →
Benchmarks
Assembly101
EPIC-KITCHENS-100
Most implemented
Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition
Opening the Vocabulary of Egocentric Actions
Papers
Robust Zero-Shot Generalization for Open-Vocabulary Action Recognition via Task Arithmetic
Open Vocabulary Action Recognition (OVAR) enables the recognition of novel actions by leveraging vision-language representations, overcoming the limitations of traditional closed-set approaches. However, achieving robust…
Open Vocabulary Action RecognitionZero-shot GeneralizationLearning to Generalize without Bias for Open-Vocabulary Action Recognition
Leveraging the effective visual-text alignment and static generalizability from CLIP, recent video learners adopt CLIP initialization with further regularization or recombination for generalization in open-vocabulary act…
Action RecognitionMeta-LearningOpen Vocabulary Action RecognitionDENOISER: Rethinking the Robustness for Open-Vocabulary Action Recognition
As one of the fundamental video tasks in computer vision, Open-Vocabulary Action Recognition (OVAR) recently gains increasing attention, with the development of vision-language pre-trainings. To enable generalization of …
Action RecognitionDenoisingOpen Vocabulary Action RecognitionRethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition
Building upon the impressive success of CLIP (Contrastive Language-Image Pretraining), recent pioneer works have proposed to adapt the powerful CLIP to video data, leading to efficient and effective video learners for op…
Action RecognitionOpen Vocabulary Action RecognitionFROSTER: Frozen CLIP Is A Strong Teacher for Open-Vocabulary Action Recognition
In this paper, we introduce FROSTER, an effective framework for open-vocabulary action recognition. The CLIP model has achieved remarkable success in a range of image-based tasks, benefiting from its strong generalizatio…
Action RecognitionOpen Vocabulary Action RecognitionOpening the Vocabulary of Egocentric Actions
Human actions in egocentric videos are often hand-object interactions composed from a verb (performed by the hand) applied to an object. Despite their extensive scaling up, egocentric datasets still face two limitations …
Action RecognitionObjectOpen Vocabulary Action Recognition