paper-with-me

홈 › Papers

VTD-CLIP: Video-to-Text Discretization via Prompting CLIP

2025-03-24 · Wencheng Zhu, Yuexin Wang, Hongxuan Li, Pengfei Zhu, Danqing Song, QinGhua Hu

Vision-language models bridge visual and linguistic understanding and have proven to be powerful for video recognition tasks. Existing approaches primarily rely on parameter-efficient fine-tuning of image-text pre-trained models, yet they often suffer from limited interpretability and poor generalization due to inadequate temporal modeling. To address these, we propose a simple yet effective video-to-text discretization framework. Our method repurposes the frozen text encoder to construct a visual codebook from video class labels due to the many-to-one contrastive alignment between visual and textual embeddings in multimodal pretraining. This codebook effectively transforms temporal visual data into textual tokens via feature lookups and offers interpretable video representations through explicit video modeling. Then, to enhance robustness against irrelevant or noisy frames, we introduce a confidence-aware fusion module that dynamically weights keyframes by assessing their semantic relevance via the codebook. Furthermore, our method incorporates learnable text prompts to conduct adaptive codebook updates. Extensive experiments on HMDB-51, UCF-101, SSv2, and Kinetics-400 have validated the superiority of our approach, achieving more competitive improvements over state-of-the-art methods. The code will be publicly available at https://github.com/isxinxin/VTD-CLIP.

📄 PDF Abstract BibTeX arXiv:2503.18407

Code (0)

등록된 구현이 없습니다.

Tasks

parameter-efficient fine-tuningVideo Recognition

Similar Papers 제목 키워드 기반

It's Just Another Day: Unique Video Captioning by Discriminative Prompting

2024-10-15 · Toby Perrett, Tengda Han, Dima Damen, Andrew Zisserman

Long videos contain many repeating actions, events and shots. These repetitions are frequently given identical captions, which makes it difficult to retrieve the exact desired clip using a text search. In this paper, we …

Video Captioning

Open-Set Video-based Facial Expression Recognition with Human Expression-sensitive Prompting

2024-04-26 · Yuanyuan Liu, Yuxuan Huang, Shuyang Liu, Yibing Zhan 외

In Video-based Facial Expression Recognition (V-FER), models are typically trained on closed-set datasets with a fixed number of known classes. However, these V-FER models cannot deal with unknown classes that are preval…

Facial Expression RecognitionMulti-Task LearningOpen Set LearningPrompt Learning+2

SignCLIP: Connecting Text and Sign Language by Contrastive Learning

2024-07-01 · Zifan Jiang, Gerard Sant, Amit Moryossef, Mathias Müller 외

We present SignCLIP, which re-purposes CLIP (Contrastive Language-Image Pretraining) to project spoken language text and sign language videos, two classes of natural languages of distinct modalities, into the same space.…

Contrastive LearningRetrievalSign Language RecognitionText Retrieval+1

Vita-CLIP: Video and text adaptive CLIP via Multimodal Prompting

2023-04-06 · CVPR 2023 1 · Syed Talal Wasim, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan 외

Adopting contrastive image-text pretrained models like CLIP towards video classification has gained attention due to its cost-effectiveness and competitive performance. However, recent works in this area face a trade-off…

Action RecognitionPrompt LearningVideo ClassificationZero-Shot Action Recognition+1

CLIP-AUTT: Test-Time Personalization with Action Unit Prompting for Fine-Grained Video Emotion Recognition

2026-03-30 · Muhammad Osama Zeeshan, Masoumeh Sharafi, Benoit Savary, Alessandro Lameiras Koerich 외 arxiv

Personalization in emotion recognition (ER) is essential for accurate interpretation of subtle and subject-specific expressive patterns. Recent advances in vision-language models (VLMs), such as CLIP, demonstrate strong …

Facial Expression RecognitionVideo Emotion Recognition