paper-with-me

Papers

Multi-grained Temporal Prototype Learning for Few-shot Video Object Segmentation

2023-09-20 · ICCV 2023 1 · Nian Liu, Kepan Nan, Wangbo Zhao, Yuanwei Liu, Xiwen Yao, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer, Junwei Han, Fahad Shahbaz Khan

Few-Shot Video Object Segmentation (FSVOS) aims to segment objects in a query video with the same category defined by a few annotated support images. However, this task was seldom explored. In this work, based on IPMT, a state-of-the-art few-shot image segmentation method that combines external support guidance information with adaptive query guidance cues, we propose to leverage multi-grained temporal guidance information for handling the temporal correlation nature of video data. We decompose the query video information into a clip prototype and a memory prototype for capturing local and long-term internal temporal guidance, respectively. Frame prototypes are further used for each frame independently to handle fine-grained adaptive guidance and enable bidirectional clip-frame prototype communication. To reduce the influence of noisy memory, we propose to leverage the structural similarity relation among different predicted regions and the support for selecting reliable memory frames. Furthermore, a new segmentation loss is also proposed to enhance the category discriminability of the learned prototypes. Experimental results demonstrate that our proposed video IPMT model significantly outperforms previous models on two benchmark datasets. Code is available at https://github.com/nankepan/VIPMT.

📄 PDF Abstract BibTeX arXiv:2309.11160

Code (1)

nankepan/VIPMT 공식 구현 pytorch

Tasks

Image SegmentationSegmentationSemantic SegmentationVideo Object SegmentationVideo Semantic Segmentation

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Compound Prototype Matching for Few-shot Action Recognition

2022-07-12 · Yifei HUANG, Lijin Yang, Yoichi Sato

Few-shot action recognition aims to recognize novel action classes using only a small number of labeled training samples. In this work, we propose a novel approach that first summarizes each video into compound prototype…

Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionVideo Similarity

Spatio-temporal Decoupled Knowledge Compensator for Few-Shot Action Recognition

2026-02-20 · Hongyu Qu, Xiangbo Shu, Rui Yan, Hailiang Gao 외 arxiv

Few-Shot Action Recognition (FSAR) is a challenging task that requires recognizing novel action categories with a few labeled videos. Recent works typically apply semantically coarse category names as auxiliary contexts …

Action Recognition

Inductive and Transductive Few-Shot Video Classification via Appearance and Temporal Alignments

2022-07-21 · Khoi D. Nguyen, Quoc-Huy Tran, Khoi Nguyen, Binh-Son Hua 외

We present a novel method for few-shot video classification, which performs appearance and temporal alignments. In particular, given a pair of query and support videos, we conduct appearance alignment via frame-level fea…

General ClassificationVideo Classification

MA-FSAR: Multimodal Adaptation of CLIP for Few-Shot Action Recognition

2023-08-03 · Jiazheng Xing, Chao Xu, Mengmeng Wang, Guang Dai 외

Applying large-scale vision-language pre-trained models like CLIP to few-shot action recognition (FSAR) can significantly enhance both performance and efficiency. While several studies have recognized this advantage, mos…

Action RecognitionFew-Shot action recognitionFew Shot Action Recognitionparameter-efficient fine-tuning

Progressive Spatio-Temporal Prototype Matching for Text-Video Retrieval

2023-01-01 · ICCV 2023 1 · Pandeng Li, Chen-Wei Xie, Liming Zhao, Hongtao Xie 외

The performance of text-video retrieval has been significantly improved by vision-language cross-modal learning schemes. The typical solution is to directly align the global video-level and sentence-level features d…

DiversityObjectRetrievalSentence+1