paper-with-me

Video Retrieval

19개 벤치마크 · 논문 584편 · 이 태스크의 논문 보기 →

Benchmarks

MSR-VTT-1kA

결과 126개

MSR-VTT

결과 81개

DiDeMo

결과 80개

LSMDC

결과 77개

ActivityNet

결과 62개

MSVD

결과 48개

FIVR-200K

결과 34개

YouCook2

결과 32개

VATEX

결과 27개

QuerYD

결과 10개

SSv2-label retrieval

결과 10개

Condensed Movies

결과 6개

EgoExoLearn

결과 4개

TGIF

결과 4개

TVR

결과 4개

Charades-STA

결과 2개

MSVD-Indonesian

결과 2개

RUDDER

결과 2개

Most implemented

Papers

Multi-modal Knowledge Preserving Adapter for Embedding Backward Compatibility

2026-09-15 · Jaeseok Byun, Gukyeong Kwon, Han-Kai Hsu, Meher Gitika Karumuri 외 arxiv

Upgrading embedding models typically requires expensive database re-indexing, as new query embeddings are incompatible with existing database embeddings. While Backward Compatible Training (BCT) mitigates this by enforci…

Video Retrieval

Beyond Similarity: Foundation Models as an Efficient Backbone for Training-Free Composed Video Retrieval

2026-09-09 · Dmitry Demidov, Muhammad Zaigham Zaheer, Omkar Thawakar, Abdelrahman Mohamed Shaker 외 arxiv

Composed video retrieval (CoVR) searches a gallery for the target video that realizes a natural-language modification of a source clip. However, at gallery scale, this creates a fundamental tension: compact embeddings en…

Video Retrieval

TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval

2026-08-13 · Yi-Chung Chen, Philip Jacobson, Tom Lampo, Yiren Lu 외 arxiv

Efficiently retrieving relevant clips from large-scale driving logs is essential for data curation, model development, and safety analysis. Structured and rule-based retrieval systems can explicitly target driving events…

Video Retrieval

Distribution-Alignment Bridge for Uncertainty-Aware Text-to-Video Retrieval

2026-07-23 · Kyeongmo Chae, Jihoon Lee, Sangtae Ahn arxiv

This paper proposes the Distribution-Alignment Bridge (DAB), a framework that reconceptualizes text-to-video retrieval as a distribution alignment task rather than traditional deterministic point matching. By modeling bo…

Video Retrieval

VEGAS: Human-Aligned Video Caption Evaluation via Gaze

2026-07-09 · Shenghui Chen, Po-han Li, Ximeng Sun, Shijia Yang 외 arxiv

Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention. We propose VEGAS (Video caption Evaluation via GAze Score), a training-free metric…

Video CaptioningVideo Retrieval

QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding

2026-07-06 · Wei Ao, Lan Wang, Vishnu Naresh Boddeti arxiv

The performance of vision-language models (VLMs) in video understanding declines with increasing video duration, as video moments unrelated to the query confuse their language components. Multimodal retrieval has emerged…

Video Retrieval

전체 584편 보기 →