Video Retrieval
19개 벤치마크 · 논문 584편 · 이 태스크의 논문 보기 →
Benchmarks
MSR-VTT-1kA
MSR-VTT
DiDeMo
LSMDC
ActivityNet
MSVD
FIVR-200K
YouCook2
VATEX
QuerYD
SSv2-label retrieval
SSv2-template retrieval
Condensed Movies
EgoExoLearn
TGIF
TVR
Charades-STA
MSVD-Indonesian
RUDDER
Most implemented
CoCa: Contrastive Captioners are Image-Text Foundation Models
ECO: Efficient Convolutional Network for Online Video Understanding
VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval
Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval
Learning a Text-Video Embedding from Incomplete and Heterogeneous Data
Papers
Multi-modal Knowledge Preserving Adapter for Embedding Backward Compatibility
Upgrading embedding models typically requires expensive database re-indexing, as new query embeddings are incompatible with existing database embeddings. While Backward Compatible Training (BCT) mitigates this by enforci…
Video RetrievalBeyond Similarity: Foundation Models as an Efficient Backbone for Training-Free Composed Video Retrieval
Composed video retrieval (CoVR) searches a gallery for the target video that realizes a natural-language modification of a source clip. However, at gallery scale, this creates a fundamental tension: compact embeddings en…
Video RetrievalTraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval
Efficiently retrieving relevant clips from large-scale driving logs is essential for data curation, model development, and safety analysis. Structured and rule-based retrieval systems can explicitly target driving events…
Video RetrievalDistribution-Alignment Bridge for Uncertainty-Aware Text-to-Video Retrieval
This paper proposes the Distribution-Alignment Bridge (DAB), a framework that reconceptualizes text-to-video retrieval as a distribution alignment task rather than traditional deterministic point matching. By modeling bo…
Video RetrievalVEGAS: Human-Aligned Video Caption Evaluation via Gaze
Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention. We propose VEGAS (Video caption Evaluation via GAze Score), a training-free metric…
Video CaptioningVideo RetrievalQSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding
The performance of vision-language models (VLMs) in video understanding declines with increasing video duration, as video moments unrelated to the query confuse their language components. Multimodal retrieval has emerged…
Video Retrieval