Papers Video Similarity
“Video Similarity” 태그가 달린 논문 24편 · 필터 해제
FiVE: A Fine-grained Video Editing Benchmark for Evaluating Emerging Diffusion and Rectified Flow Models
Numerous text-to-video (T2V) editing methods have emerged recently, but the lack of a standardized benchmark for fair evaluation has led to inconsistent claims and an inability to assess model sensitivity to hyperparamet…
SensitivityVideo EditingVideo SimilarityNarrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level Captions
In recent text-video retrieval, the use of additional captions from vision-language models has shown promising effects on the performance. However, existing models using additional captions often have struggled to captur…
RetrievalVideo RetrievalVideo SimilarityContextual Augmented Global Contrast for Multimodal Intent Recognition
Multimodal intent recognition (MIR) aims to perceive the human intent polarity via language visual and acoustic modalities. The inherent intent ambiguity makes it challenging to recognize in multimodal scenarios. Exi…
Contrastive LearningIntent RecognitionMultimodal Intent RecognitionMultimodal Sentiment Analysis+2Network-Based Video Recommendation Using Viewing Patterns and Modularity Analysis: An Integrated Framework
The proliferation of video-on-demand (VOD) services has led to a paradox of choice, overwhelming users with vast content libraries and revealing limitations in current recommender systems. This research introduces a nove…
ClusteringCollaborative FilteringRecommendation SystemsVideo SimilarityThe 2023 Video Similarity Dataset and Challenge
This work introduces a dataset, benchmark, and challenge for the problem of video copy detection and localization. The problem comprises two distinct but related tasks: determining whether a query video shares content wi…
Copy DetectionVideo SimilarityA Similarity Alignment Model for Video Copy Segment Matching
With the development of multimedia technology, Video Copy Detection has been a crucial problem for social media platforms. Meta AI hold Video Similarity Challenge on CVPR 2023 to push the technology forward. In this repo…
Copy DetectionPartial Video Copy DetectionVideo SimilarityA Dual-level Detection Method for Video Copy Detection
With the development of multimedia technology, Video Copy Detection has been a crucial problem for social media platforms. Meta AI hold Video Similarity Challenge on CVPR 2023 to push the technology forward. In this pape…
Copy DetectionPartial Video Copy DetectionVideo EditingVideo SimilarityFew-shot Action Recognition via Intra- and Inter-Video Information Maximization
Current few-shot action recognition involves two primary sources of information for classification:(1) intra-video information, determined by frame content within a single video clip, and (2) inter-video information, mea…
Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionTemporal Action Localization+13rd Place Solution to Meta AI Video Similarity Challenge
This paper presents our 3rd place solution in both Descriptor Track and Matching Track of the Meta AI Video Similarity Challenge (VSC2022), a competition aimed at detecting video copies. Our approach builds upon existing…
Copy DetectionVideo SimilarityFeature-compatible Progressive Learning for Video Copy Detection
Video Copy Detection (VCD) has been developed to identify instances of unauthorized or duplicated video content. This paper presents our second place solutions to the Meta AI Video Similarity Challenge (VSC22), CVPR 2023…
Copy DetectionVideo SimilaritySelf-Supervised Video Similarity Learning
We introduce S$^2$VS, a video similarity learning approach with self-supervision. Self-Supervised Learning (SSL) is typically used to train deep models on a proxy task so as to have strong transferability on target tasks…
ISVRRetrievalSelf-Supervised LearningVideo Retrieval+2Contrastive Masked Autoencoders for Self-Supervised Video Hashing
Self-Supervised Video Hashing (SSVH) models learn to generate short binary representations for videos without ground-truth supervision, facilitating large-scale video retrieval efficiency and attracting increasing resear…
DecoderRetrievalVideo RetrievalVideo Similarity+13D-CSL: self-supervised 3D context similarity learning for Near-Duplicate Video Retrieval
In this paper, we introduce 3D-CSL, a compact pipeline for Near-Duplicate Video Retrieval (NDVR), and explore a novel self-supervised learning strategy for video similarity learning. Most previous methods only extract vi…
RetrievalSelf-Supervised LearningTripletVideo Prediction+2C2KD: Cross-Lingual Cross-Modal Knowledge Distillation for Multilingual Text-Video Retrieval
Multilingual text-video retrieval methods have improved significantly in recent years, but the performance for other languages lags behind English. We propose a Cross-Lingual Cross-Modal Knowledge Distillation method to …
Knowledge DistillationRetrievalVideo RetrievalVideo SimilarityCompound Prototype Matching for Few-shot Action Recognition
Few-shot action recognition aims to recognize novel action classes using only a small number of labeled training samples. In this work, we propose a novel approach that first summarizes each video into compound prototype…
Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionVideo SimilarityCross-Lingual Cross-Modal Consolidation for Effective Multilingual Video Corpus Moment Retrieval
Existing multilingual video corpus moment retrieval (mVCMR) methods are mainly based on a two-stream structure. The visual stream utilizes the visual content in the video to estimate the query-visual similarity, and the …
Moment RetrievalRetrievalVideo Corpus Moment RetrievalVideo SimilarityTencent-MVSE: A Large-Scale Benchmark Dataset for Multi-Modal Video Similarity Evaluation
Multi-modal video similarity evaluation is important for video recommendation systems such as video de-duplication, relevance matching, ranking, and diversity control. However, there still lacks a benchmark dataset t…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DiversityRecommendation Systems+3Top1 Solution of QQ Browser 2021 Ai Algorithm Competition Track 1 : Multimodal Video Similarity
In this paper, we describe the solution to the QQ Browser 2021 Ai Algorithm Competition (AIAC) Track 1. We use the multi-modal transformer model for the video embedding extraction. In the pretrain phase, we train the mod…
Language ModelingLanguage ModellingTAGVideo SimilarityBiC-Net: Learning Efficient Spatio-Temporal Relation for Text-Video Retrieval
The task of text-video retrieval aims to understand the correspondence between language and vision, has gained increasing attention in recent years. Previous studies either adopt off-the-shelf 2D/3D-CNN and then use aver…
Cross-Modal RetrievalRelationRetrievalVideo Retrieval+1Video Similarity and Alignment Learning on Partial Video Copy Detection
Existing video copy detection methods generally measure video similarity based on spatial similarities between key frames, neglecting the latent similarity in temporal dimension, so that the video similarity is biased to…
Copy DetectionPartial Video Copy DetectionVideo Similarity