paper-with-me

Papers

ZEETAD: Adapting Pretrained Vision-Language Model for Zero-Shot End-to-End Temporal Action Detection

2023-11-01 · Thinh Phan, Khoa Vo, Duy Le, Gianfranco Doretto, Donald Adjeroh, Ngan Le

Temporal action detection (TAD) involves the localization and classification of action instances within untrimmed videos. While standard TAD follows fully supervised learning with closed-set setting on large training data, recent zero-shot TAD methods showcase the promising open-set setting by leveraging large-scale contrastive visual-language (ViL) pretrained models. However, existing zero-shot TAD methods have limitations on how to properly construct the strong relationship between two interdependent tasks of localization and classification and adapt ViL model to video understanding. In this work, we present ZEETAD, featuring two modules: dual-localization and zero-shot proposal classification. The former is a Transformer-based module that detects action events while selectively collecting crucial semantic embeddings for later recognition. The latter one, CLIP-based module, generates semantic embeddings from text and frame inputs for each temporal unit. Additionally, we enhance discriminative capability on unseen classes by minimally updating the frozen CLIP encoder with lightweight adapters. Extensive experiments on THUMOS14 and ActivityNet-1.3 datasets demonstrate our approach's superior performance in zero-shot TAD and effective knowledge transfer from ViL models to unseen action categories.

📄 PDF Abstract BibTeX arXiv:2311.00729

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionClassificationLanguage ModelingLanguage ModellingTransfer LearningVideo Understanding

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Zero-Shot and Few-Shot Video Question Answering with Multi-Modal Prompts

2023-09-27 · Deniz Engin, Yannis Avrithis

Recent vision-language models are driven by large-scale pretrained models. However, adapting pretrained models on limited data presents challenges such as overfitting, catastrophic forgetting, and the cross-modal gap bet…

Few-shot Video Question AnsweringPrompt LearningQuestion AnsweringVideo Question Answering+1

Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models

2025-06-10 · Chenyu Lian, Hong-Yu Zhou, Dongyun Liang, Jing Qin 외

Medical vision-language alignment through cross-modal contrastive learning shows promising performance in image-text matching tasks, such as retrieval and zero-shot classification. However, conventional cross-modal contr…

Contrastive LearningImage-text matchingImage to textImage-to-Text Retrieval+5

VLM-AutoDrive: Post-Training Vision-Language Models for Safety-Critical Autonomous Driving Events

2026-03-18 · Mohammad Qazim Bhat, Yufan Huang, Niket Agarwal, Hao Wang 외 arxiv

The rapid growth of ego-centric dashcam footage presents a major challenge for detecting safety-critical events such as collisions and near-collisions, scenarios that are brief, rare, and difficult for generic vision mod…

Visual Question AnsweringAutonomous DrivingAnomaly Detection

LiT Tuned Models for Efficient Species Detection

2023-02-12 · Andre Nakkab, Benjamin Feuer, Chinmay Hegde

Recent advances in training vision-language models have demonstrated unprecedented robustness and transfer learning effectiveness; however, standard computer vision datasets are image-only, and therefore not well adapted…

Fine-Grained Image Classificationimage-classificationImage ClassificationTransfer Learning+2

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images

2025-08-29 · Juneyoung Ro, Namwoo Kim, Yoonjin Yoon arxiv

Effectively understanding urban scenes requires fine-grained spatial reasoning about objects, layouts, and depth cues. However, how well current vision-language models (VLMs), pretrained on general scenes, transfer these…

Spatial ReasoningObject Detection