Papers Multi-Instance Retrieval
“Multi-Instance Retrieval” 태그가 달린 논문 19편 · 필터 해제
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization
Egocentric video-language understanding demands both high efficiency and accurate spatial-temporal modeling. Existing approaches face three key challenges: 1) Excessive pre-training cost arising from multi-stage pre-trai…
Multi-Instance RetrievalRetrievalVideo UnderstandingContextRefine-CLIP for EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2025
This report presents ContextRefine-CLIP (CR-CLIP), an efficient model for visual-textual multi-instance retrieval tasks. The approach is based on the dual-encoder AVION, on which we introduce a cross-modal attention flow…
Cross-Modal RetrievalEnsemble LearningMulti-Instance RetrievalRetrievalModeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning
In egocentric video understanding, the motion of hands and objects as well as their interactions play a significant role by nature. However, existing egocentric video representation learning methods mainly focus on align…
Large Language ModelMulti-Instance RetrievalObjectRepresentation Learning+2Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning
We present EMBED (Egocentric Models Built with Exocentric Data), a method designed to transform exocentric video-language data for egocentric video representation learning. Large-scale exocentric data covers diverse acti…
Multi-Instance RetrievalRepresentation LearningStyle TransferEgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation
In this report, we present our solutions to the EgoVis Challenges in CVPR 2024, including five tracks in the Ego4D challenge and three tracks in the EPIC-Kitchens challenge. Building upon the video-language two-tower mod…
Action AnticipationAction RecognitionDomain AdaptationLong Term Action Anticipation+4Symmetric Multi-Similarity Loss for EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2024
In this report, we present our champion solution for EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge in CVPR 2024. Essentially, this challenge differs from traditional visual-text retrieval tasks by providing a corr…
Ensemble LearningMulti-Instance RetrievalRetrievalText RetrievalEgoNCE++: Do Egocentric Video-Language Models Really Understand Hand-Object Interactions?
Egocentric video-language pretraining is a crucial paradigm to advance the learning of egocentric hand-object interactions (EgoHOI). Despite the great success on existing testbeds, these benchmarks focus more on closed-s…
Action RecognitionAttributeIn-Context LearningMulti-Instance RetrievalTraining a Large Video Model on a Single Machine in a Day
Videos are big, complex to pre-process, and slow to train on. State-of-the-art large-scale video models are trained on clusters of 32 or more GPUs for several days. As a consequence, academia largely ceded the training o…
Action RecognitionCPUGPUMulti-Instance RetrievalEgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone
Video-language pre-training (VLP) has become increasingly important due to its ability to generalize to various vision and language tasks. However, existing egocentric VLP frameworks utilize separate video and language e…
Action RecognitionMoment QueriesMulti-Instance RetrievalNatural Language Queries+2UniUD Submission to the EPIC-Kitchens-100 Multi-Instance Retrieval Challenge 2023
In this report, we present the technical details of our submission to the EPIC-Kitchens-100 Multi-Instance Retrieval Challenge 2023. To participate in the challenge, we ensembled two models trained with two different los…
Multi-Instance RetrievalRetrievalHierVL: Learning Hierarchical Video-Language Embeddings
Video-language embeddings are a promising avenue for injecting semantics into visual representations, but existing methods capture only short-term associations between seconds-long video clips and their accompanying text…
Action ClassificationAction RecognitionLong Term Action AnticipationLong Term Anticipation+1Learning Video Representations from Large Language Models
We introduce LaViLa, a new approach to learning video-language representations by leveraging Large Language Models (LLMs). We repurpose pre-trained LLMs to be conditioned on visual input, and finetune them to create auto…
Action ClassificationAction RecognitionDiversityEgocentric Activity Recognition+2Egocentric Video-Language Pretraining @ EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2022
In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for the EPIC-KITCHENS-100 Multi-Instance Retrieval (MIR) challenge. Especially, we exploit the recently released Ego4D d…
Language ModelingLanguage ModellingMulti-Instance RetrievalRetrievalExploiting Semantic Role Contextualized Video Features for Multi-Instance Text-Video Retrieval EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2022
In this report, we present our approach for EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2022. We first parse sentences into semantic roles corresponding to verbs and nouns; then utilize self-attentions to exploi…
Multi-Instance RetrievalRetrievalSemantic SimilaritySemantic Textual Similarity+2UniUD-FBK-UB-UniBZ Submission to the EPIC-Kitchens-100 Multi-Instance Retrieval Challenge 2022
This report presents the technical details of our submission to the EPIC-Kitchens-100 Multi-Instance Retrieval Challenge 2022. To participate in the challenge, we designed an ensemble consisting of different models train…
Multi-Instance RetrievalRetrievalTripletEgocentric Video-Language Pretraining
Video-Language Pretraining (VLP), which aims to learn transferable representation to advance a wide range of video-text downstream tasks, has recently received increasing attention. Best performing works rely on large-sc…
Action RecognitionContrastive LearningMoment QueriesMulti-Instance Retrieval+9Relevance-based Margin for Contrastively-trained Video Retrieval Models
Video retrieval using natural language queries has attracted increasing interest due to its relevance in real-world applications, from intelligent access in private media galleries to web-scale video search. Learning the…
Multi-Instance RetrievalNatural Language QueriesRetrievalVideo RetrievalLearning video retrieval models with relevance-aware online mining
Due to the amount of videos and related captions uploaded every hour, deep learning-based solutions for cross-modal video retrieval are attracting more and more attention. A typical approach consists in learning a joint …
Multi-Instance RetrievalRetrievalvalidVideo RetrievalStochastic Learning of Multi-Instance Dictionary for Earth Mover's Distance based Histogram Comparison
Dictionary plays an important role in multi-instance data representation. It maps bags of instances to histograms. Earth mover's distance (EMD) is the most effective histogram distance metric for the application of multi…
Dictionary LearningImage RetrievalMedical Image RetrievalMulti-Instance Retrieval+4