paper-with-me

Open Vocabulary Action Recognition

2개 벤치마크 · 논문 6편 · 이 태스크의 논문 보기 →

Benchmarks

Assembly101

결과 2개

EPIC-KITCHENS-100

결과 2개

Most implemented

Papers

Robust Zero-Shot Generalization for Open-Vocabulary Action Recognition via Task Arithmetic

2026-06-17 · Francesca Morandi, Omayma Moussadek, Federico Venturini, Mauro Suardi 외 arxiv

Open Vocabulary Action Recognition (OVAR) enables the recognition of novel actions by leveraging vision-language representations, overcoming the limitations of traditional closed-set approaches. However, achieving robust…

Open Vocabulary Action RecognitionZero-shot Generalization

Learning to Generalize without Bias for Open-Vocabulary Action Recognition

2025-02-27 · Yating Yu, Congqi Cao, Yifan Zhang, Yanning Zhang

Leveraging the effective visual-text alignment and static generalizability from CLIP, recent video learners adopt CLIP initialization with further regularization or recombination for generalization in open-vocabulary act…

Action RecognitionMeta-LearningOpen Vocabulary Action Recognition

DENOISER: Rethinking the Robustness for Open-Vocabulary Action Recognition

2024-04-23 · Haozhe Cheng, Cheng Ju, Haicheng Wang, Jinxiang Liu 외

As one of the fundamental video tasks in computer vision, Open-Vocabulary Action Recognition (OVAR) recently gains increasing attention, with the development of vision-language pre-trainings. To enable generalization of …

Action RecognitionDenoisingOpen Vocabulary Action Recognition

Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition

2024-03-03 · Kun-Yu Lin, Henghui Ding, Jiaming Zhou, Yu-Ming Tang 외

Building upon the impressive success of CLIP (Contrastive Language-Image Pretraining), recent pioneer works have proposed to adapt the powerful CLIP to video data, leading to efficient and effective video learners for op…

Action RecognitionOpen Vocabulary Action Recognition

FROSTER: Frozen CLIP Is A Strong Teacher for Open-Vocabulary Action Recognition

2024-02-05 · Xiaohu Huang, Hao Zhou, Kun Yao, Kai Han

In this paper, we introduce FROSTER, an effective framework for open-vocabulary action recognition. The CLIP model has achieved remarkable success in a range of image-based tasks, benefiting from its strong generalizatio…

Action RecognitionOpen Vocabulary Action Recognition

Opening the Vocabulary of Egocentric Actions

2023-08-22 · NeurIPS 2023 11 · Dibyadip Chatterjee, Fadime Sener, Shugao Ma, Angela Yao

Human actions in egocentric videos are often hand-object interactions composed from a verb (performed by the hand) applied to an object. Despite their extensive scaling up, egocentric datasets still face two limitations …

Action RecognitionObjectOpen Vocabulary Action Recognition