paper-with-me

Papers

Robust Zero-Shot Generalization for Open-Vocabulary Action Recognition via Task Arithmetic

2026-06-17 · Francesca Morandi, Omayma Moussadek, Federico Venturini, Mauro Suardi, Alessandro Banzatti, Francesco Cannarile, Angelo Porrello, Simone Calderara arxiv

Open Vocabulary Action Recognition (OVAR) enables the recognition of novel actions by leveraging vision-language representations, overcoming the limitations of traditional closed-set approaches. However, achieving robust performance in real-world scenarios typically requires domain-specific fine-tuning, which is often costly and raises privacy and regulatory concerns. In this work, we propose an alternative paradigm that bypasses target-domain training and recombines knowledge from existing datasets and models. Leveraging model merging and task arithmetic, we extract and combine task vectors from models fine-tuned on diverse public OVAR datasets. We show that, in out-of-distribution settings, the resulting merged model achieves superior zero-shot generalization to the pre-trained base model. Code is available at https://github.com/omaymaMoussadek/robust-ovar

📄 PDF Abstract BibTeX arXiv:2606.20734

Code (0)

등록된 구현이 없습니다.

Tasks

Open Vocabulary Action RecognitionZero-shot Generalization

Similar Papers 제목 키워드 기반

Open-Vocabulary Audio-Visual Semantic Segmentation

2024-07-31

Audio-visual semantic segmentation (AVSS) aims to segment and classify sounding objects in videos with acoustic cues. However, most approaches operate on the close-set assumption and only identify pre-defined categories …

Open-VCLIP: Transforming CLIP to an Open-vocabulary Video Model via Interpolated Weight Optimization

2023-02-01 · Zejia Weng, Xitong Yang, Ang Li, Zuxuan Wu 외

Contrastive Language-Image Pretraining (CLIP) has demonstrated impressive zero-shot learning abilities for image understanding, yet limited effort has been made to investigate CLIP for zero-shot video recognition. We int…

Action RecognitionContinual LearningVideo RecognitionZero-Shot Learning

Boosting Segment Anything Model Towards Open-Vocabulary Learning

2023-12-06 · Xumeng Han, Longhui Wei, Xuehui Yu, Zhiyang Dou 외

The recent Segment Anything Model (SAM) has emerged as a new paradigmatic vision foundation model, showcasing potent zero-shot generalization and flexible prompting. Despite SAM finding applications and adaptations in va…

modelObjectObject LocalizationRegion Proposal+1

Generating Action-conditioned Prompts for Open-vocabulary Video Action Recognition

2023-12-04 · Chengyou Jia, Minnan Luo, Xiaojun Chang, Zhuohang Dang 외

Exploring open-vocabulary video action recognition is a promising venture, which aims to recognize previously unseen actions within any arbitrary set of categories. Existing methods typically adapt pretrained image-text …

Action RecognitionDescriptiveTemporal Action Localization

Exploring Vision-Language Models for Open-Vocabulary Zero-Shot Action Segmentation

2026-02-24 · Asim Unmesh, Kaki Ramesh, Mayank Patel, Rahul Jain 외 arxiv

Temporal Action Segmentation (TAS) requires dividing videos into action segments, yet the vast space of activities and alternative breakdowns makes collecting comprehensive datasets infeasible. Existing methods remain li…

Action Segmentation