paper-with-me

Video Retrieval 벤치마크

Video Retrieval on VATEX

27개 결과 · ⬇ CSV · JSON

text-to-video R@1

57.3 67.88 78.45 89.02 99.6 2021-06 2026-09 CLIP2Video — 57.3 (2021-06-21) CLIP2Video — 57.3 (2021-06-21) LAFF — 59.1 (2021-12-03) LAFF — 59.1 (2021-12-03) QB-Norm+CLIP2Video — 58.8 (2021-12-23) QB-Norm+CLIP2Video — 58.8 (2021-12-23) TS2-Net — 59.1 (2022-07-16) TS2-Net — 59.1 (2022-07-16) InternVideo — 71.1 (2022-12-06) InternVideo — 71.1 (2022-12-06) Cap4Video — 66.6 (2022-12-31) Cap4Video — 66.6 (2022-12-31) Unmasked Teacher — 72.0 (2023-03-28) Unmasked Teacher — 72.0 (2023-03-28) VALOR — 78.5 (2023-04-17) VALOR — 78.5 (2023-04-17) VAST — 83.0 (2023-05-29) VAST — 83.0 (2023-05-29) Side4Video — 68.8 (2023-11-27) Side4Video — 68.8 (2023-11-27) TeachCLIP — 63.6 (2024-01-01) TeachCLIP — 63.6 (2024-01-01) InternVideo2-6B — 75.5 (2024-03-22) InternVideo2-6B — 75.5 (2024-03-22) GRAM — 87.7 (2024-12-16) GRAM — 87.7 (2024-12-16) Vote-in-Context — 99.6 (2025-11-03) CLIP2Video — 57.3 (2021-06-21) LAFF — 59.1 (2021-12-03) InternVideo — 71.1 (2022-12-06) Unmasked Teacher — 72.0 (2023-03-28) VALOR — 78.5 (2023-04-17) VAST — 83.0 (2023-05-29) GRAM — 87.7 (2024-12-16) Vote-in-Context — 99.6 (2025-11-03)
RankModel text-to-video R@1text-to-video R@5text-to-video R@10text-to-video R@50text-to-video MedianRtext-to-video MeanRvideo-to-text R@1video-to-text R@10 Extra Training Data PaperCodeYear
1 Vote-in-Context 자동 추출 99.6 Vote-in-Context: Turning VLMs into Zero-Shot Rank Fusers 2025
2 GRAM 87.710084.6100 Gramian Multimodal Representation Learning and Alignment ispamm/GRAM · luigisigillo/gwit 2024
3 VAST 83.098.299.2 VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset TXH-mercury/VALOR · txh-mercury/vast 2023
4 VALOR 78.597.198.7 VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset TXH-mercury/VALOR 2023
5 InternVideo2-6B 75.589.3 InternVideo2: Scaling Foundation Models for Multimodal Video Understanding opengvlab/internvideo · opengvlab/internvideo2 2024
6 Unmasked Teacher 7295.197.886.099.6 Unmasked Teacher: Towards Training-Efficient Video Foundation Models opengvlab/unmasked_teacher 2023
7 InternVideo 71.187.2 InternVideo: General Video Foundation Models via Generative and Discriminative Learning opengvlab/internvideo · yingsen1/unimd 2022
8 Side4Video 68.893.597.01.02.7 Side4Video: Spatial-Temporal Side Network for Memory-Efficient Image-to-Video Transfer Learning whwu95/ATM · HJYao00/Side4Video 2023
9 Cap4Video 66.693.197.012.780.999.6 Cap4Video: What Can Auxiliary Captions Do for Text-Video Retrieval? whwu95/Cap4Video · whwu95/text4vis · whwu95/GPT4Vis · +1 2022
10 TeachCLIP 63.691.996.1 Holistic Features are almost Sufficient for Text-to-Video Retrieval ruc-aimc-lab/teachclip 2024
11 TS2-Net 59.195.2 TS2-Net: Token Shift and Selection Transformer for Text-Video Retrieval yuqi657/ts2_net 2022
11 LAFF 59.191.796.3 Lightweight Attentional Feature Fusion: A New Baseline for Text-to-Video Retrieval ruc-aimc-lab/laff 2021
13 QB-Norm+CLIP2Video 58.893.8 Cross Modal Retrieval with Querybank Normalisation ioanacroi/qb-norm 2021
14 CLIP2Video 57.39095.5 CLIP2Video: Mastering Video-Text Retrieval via Image CLIP CryhanFang/CLIP2Video 2021
15 GRAM 87.710084.6100 Gramian Multimodal Representation Learning and Alignment ispamm/GRAM · luigisigillo/gwit 2024
16 VAST 83.098.299.2 VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset TXH-mercury/VALOR · txh-mercury/vast 2023
17 VALOR 78.597.198.7 VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset TXH-mercury/VALOR 2023
18 InternVideo2-6B 75.589.3 InternVideo2: Scaling Foundation Models for Multimodal Video Understanding opengvlab/internvideo · opengvlab/internvideo2 2024
19 Unmasked Teacher 7295.197.886.099.6 Unmasked Teacher: Towards Training-Efficient Video Foundation Models opengvlab/unmasked_teacher 2023
20 InternVideo 71.187.2 InternVideo: General Video Foundation Models via Generative and Discriminative Learning opengvlab/internvideo · yingsen1/unimd 2022
1–20 / 27 다음 → 페이지당 10 20 50 100