paper-with-me

홈 › Papers

Self-supervised Video Retrieval Transformer Network

2021-04-16 · Xiangteng He, Yulin Pan, Mingqian Tang, Yiliang Lv

Content-based video retrieval aims to find videos from a large video database that are similar to or even near-duplicate of a given query video. Video representation and similarity search algorithms are crucial to any video retrieval system. To derive effective video representation, most video retrieval systems require a large amount of manually annotated data for training, making it costly inefficient. In addition, most retrieval systems are based on frame-level features for video similarity searching, making it expensive both storage wise and search wise. We propose a novel video retrieval system, termed SVRTN, that effectively addresses the above shortcomings. It first applies self-supervised training to effectively learn video representation from unlabeled data to avoid the expensive cost of manual annotation. Then, it exploits transformer structure to aggregate frame-level features into clip-level to reduce both storage space and search complexity. It can learn the complementary and discriminative information from the interactions among clip frames, as well as acquire the frame permutation and missing invariant ability to support more flexible retrieval manners. Comprehensive experiments on two challenging video retrieval datasets, namely FIVR-200K and SVD, verify the effectiveness of our proposed SVRTN method, which achieves the best performance of video retrieval on accuracy and efficiency.

📄 PDF Abstract BibTeX arXiv:2104.07993

Code (0)

등록된 구현이 없습니다.

Tasks

RetrievalSelf-supervised Video RetrievalVideo RetrievalVideo Similarity

Similar Papers 제목 키워드 기반

3D-CSL: self-supervised 3D context similarity learning for Near-Duplicate Video Retrieval

2022-11-10 · Rui Deng, Qian Wu, Yuke Li

In this paper, we introduce 3D-CSL, a compact pipeline for Near-Duplicate Video Retrieval (NDVR), and explore a novel self-supervised learning strategy for video similarity learning. Most previous methods only extract vi…

RetrievalSelf-Supervised LearningTripletVideo Prediction+2

Self-Supervised Video Hashing via Bidirectional Transformers

2021-06-19 · CVPR 2021 1 · Shuyan Li, Xiu Li, Jiwen Lu, Jie zhou

Most existing unsupervised video hashing methods are built on unidirectional models with less reliable training objectives, which underuse the correlations among frames and the similarity structure between videos. To…

DecoderRetrievalVideo Retrieval

Cross-Architecture Self-supervised Video Representation Learning

2022-05-26 · CVPR 2022 1 · Sheng Guo, Zihua Xiong, Yujie Zhong, LiMin Wang 외

In this paper, we present a new cross-architecture contrastive learning (CACL) framework for self-supervised video representation learning. CACL consists of a 3D CNN and a video transformer which are used in parallel to …

Action RecognitionContrastive LearningRepresentation LearningRetrieval+2

Self-supervised and Weakly Supervised Contrastive Learning for Frame-wise Action Representations

2022-12-06 · Minghao Chen, Renbo Tu, Chenxi Huang, Yuqi Lin 외

Previous work on action representation learning focused on global representations for short video clips. In contrast, many practical applications, such as video alignment, strongly demand learning the intensive represent…

Action ClassificationContrastive LearningRepresentation LearningRetrieval+2

Less than Few: Self-Shot Video Instance Segmentation

2022-04-19 · Pengwan Yang, Yuki M. Asano, Pascal Mettes, Cees G. M. Snoek

The goal of this paper is to bypass the need for labelled examples in few-shot video understanding at run time. While proven effective, in many practical video settings even labelling a few examples appears unrealistic. …

Few-Shot LearningInstance SegmentationRetrievalSelf-Supervised Learning+3