paper-with-me

Papers Multi-Instance Retrieval

“Multi-Instance Retrieval” 태그가 달린 논문 19편 · 필터 해제

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization

2025-06-17 · Xiaoqi Wang, Yi Wang, Lap-Pui Chau

Egocentric video-language understanding demands both high efficiency and accurate spatial-temporal modeling. Existing approaches face three key challenges: 1) Excessive pre-training cost arising from multi-stage pre-trai…

Multi-Instance RetrievalRetrievalVideo Understanding

ContextRefine-CLIP for EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2025

2025-06-12 · Jing He, YiQing Wang, Lingling Li, Kexin Zhang 외

This report presents ContextRefine-CLIP (CR-CLIP), an efficient model for visual-textual multi-instance retrieval tasks. The approach is based on the dual-encoder AVION, on which we introduce a cross-modal attention flow…

Cross-Modal RetrievalEnsemble LearningMulti-Instance RetrievalRetrieval

Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning

2025-03-02 · Baoqi Pei, Yifei HUANG, Jilan Xu, Guo Chen 외

In egocentric video understanding, the motion of hands and objects as well as their interactions play a significant role by nature. However, existing egocentric video representation learning methods mainly focus on align…

Large Language ModelMulti-Instance RetrievalObjectRepresentation Learning+2

Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning

2024-08-07 · Zi-Yi Dou, Xitong Yang, Tushar Nagarajan, Huiyu Wang 외

We present EMBED (Egocentric Models Built with Exocentric Data), a method designed to transform exocentric video-language data for egocentric video representation learning. Large-scale exocentric data covers diverse acti…

Multi-Instance RetrievalRepresentation LearningStyle Transfer

EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation

2024-06-26 · Baoqi Pei, Guo Chen, Jilan Xu, Yuping He 외

In this report, we present our solutions to the EgoVis Challenges in CVPR 2024, including five tracks in the Ego4D challenge and three tracks in the EPIC-Kitchens challenge. Building upon the video-language two-tower mod…

Action AnticipationAction RecognitionDomain AdaptationLong Term Action Anticipation+4

Symmetric Multi-Similarity Loss for EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2024

2024-06-18 · Xiaoqi Wang, Yi Wang, Lap-Pui Chau

In this report, we present our champion solution for EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge in CVPR 2024. Essentially, this challenge differs from traditional visual-text retrieval tasks by providing a corr…

Ensemble LearningMulti-Instance RetrievalRetrievalText Retrieval

EgoNCE++: Do Egocentric Video-Language Models Really Understand Hand-Object Interactions?

2024-05-28 · Boshen Xu, Ziheng Wang, Yang Du, Zhinan Song 외

Egocentric video-language pretraining is a crucial paradigm to advance the learning of egocentric hand-object interactions (EgoHOI). Despite the great success on existing testbeds, these benchmarks focus more on closed-s…

Action RecognitionAttributeIn-Context LearningMulti-Instance Retrieval

Training a Large Video Model on a Single Machine in a Day

2023-09-28 · Yue Zhao, Philipp Krähenbühl

Videos are big, complex to pre-process, and slow to train on. State-of-the-art large-scale video models are trained on clusters of 32 or more GPUs for several days. As a consequence, academia largely ceded the training o…

Action RecognitionCPUGPUMulti-Instance Retrieval

EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone

2023-07-11 · ICCV 2023 1 · Shraman Pramanick, Yale Song, Sayan Nag, Kevin Qinghong Lin 외

Video-language pre-training (VLP) has become increasingly important due to its ability to generalize to various vision and language tasks. However, existing egocentric VLP frameworks utilize separate video and language e…

Action RecognitionMoment QueriesMulti-Instance RetrievalNatural Language Queries+2

UniUD Submission to the EPIC-Kitchens-100 Multi-Instance Retrieval Challenge 2023

2023-06-27 · Alex Falcon, Giuseppe Serra

In this report, we present the technical details of our submission to the EPIC-Kitchens-100 Multi-Instance Retrieval Challenge 2023. To participate in the challenge, we ensembled two models trained with two different los…

Multi-Instance RetrievalRetrieval

HierVL: Learning Hierarchical Video-Language Embeddings

2023-01-05 · CVPR 2023 1 · Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, Kristen Grauman

Video-language embeddings are a promising avenue for injecting semantics into visual representations, but existing methods capture only short-term associations between seconds-long video clips and their accompanying text…

Action ClassificationAction RecognitionLong Term Action AnticipationLong Term Anticipation+1

Learning Video Representations from Large Language Models

2022-12-08 · CVPR 2023 1 · Yue Zhao, Ishan Misra, Philipp Krähenbühl, Rohit Girdhar

We introduce LaViLa, a new approach to learning video-language representations by leveraging Large Language Models (LLMs). We repurpose pre-trained LLMs to be conditioned on visual input, and finetune them to create auto…

Action ClassificationAction RecognitionDiversityEgocentric Activity Recognition+2

Egocentric Video-Language Pretraining @ EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2022

2022-07-04 · Kevin Qinghong Lin, Alex Jinpeng Wang, Rui Yan, Eric Zhongcong Xu 외

In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for the EPIC-KITCHENS-100 Multi-Instance Retrieval (MIR) challenge. Especially, we exploit the recently released Ego4D d…

Language ModelingLanguage ModellingMulti-Instance RetrievalRetrieval

Exploiting Semantic Role Contextualized Video Features for Multi-Instance Text-Video Retrieval EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2022

2022-06-29 · Burak Satar, Hongyuan Zhu, Hanwang Zhang, Joo Hwee Lim

In this report, we present our approach for EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2022. We first parse sentences into semantic roles corresponding to verbs and nouns; then utilize self-attentions to exploi…

Multi-Instance RetrievalRetrievalSemantic SimilaritySemantic Textual Similarity+2

UniUD-FBK-UB-UniBZ Submission to the EPIC-Kitchens-100 Multi-Instance Retrieval Challenge 2022

2022-06-22 · Alex Falcon, Giuseppe Serra, Sergio Escalera, Oswald Lanz

This report presents the technical details of our submission to the EPIC-Kitchens-100 Multi-Instance Retrieval Challenge 2022. To participate in the challenge, we designed an ensemble consisting of different models train…

Multi-Instance RetrievalRetrievalTriplet

Egocentric Video-Language Pretraining

2022-06-03 · Kevin Qinghong Lin, Alex Jinpeng Wang, Mattia Soldan, Michael Wray 외

Video-Language Pretraining (VLP), which aims to learn transferable representation to advance a wide range of video-text downstream tasks, has recently received increasing attention. Best performing works rely on large-sc…

Action RecognitionContrastive LearningMoment QueriesMulti-Instance Retrieval+9

Relevance-based Margin for Contrastively-trained Video Retrieval Models

2022-04-27 · Alex Falcon, Swathikiran Sudhakaran, Giuseppe Serra, Sergio Escalera 외

Video retrieval using natural language queries has attracted increasing interest due to its relevance in real-world applications, from intelligent access in private media galleries to web-scale video search. Learning the…

Multi-Instance RetrievalNatural Language QueriesRetrievalVideo Retrieval

Learning video retrieval models with relevance-aware online mining

2022-03-16 · Alex Falcon, Giuseppe Serra, Oswald Lanz

Due to the amount of videos and related captions uploaded every hour, deep learning-based solutions for cross-modal video retrieval are attracting more and more attention. A typical approach consists in learning a joint …

Multi-Instance RetrievalRetrievalvalidVideo Retrieval

Stochastic Learning of Multi-Instance Dictionary for Earth Mover's Distance based Histogram Comparison

2016-09-03 · Jihong Fan, Ru-Ze Liang

Dictionary plays an important role in multi-instance data representation. It maps bags of instances to histograms. Earth mover's distance (EMD) is the most effective histogram distance metric for the application of multi…

Dictionary LearningImage RetrievalMedical Image RetrievalMulti-Instance Retrieval+4
1–19 / 19