paper-with-me

Papers

Dual Encoding for Zero-Example Video Retrieval

2018-09-17 · CVPR 2019 6 · Jianfeng Dong, Xirong Li, Chaoxi Xu, Shouling Ji, Yuan He, Gang Yang, Xun Wang

This paper attacks the challenging problem of zero-example video retrieval. In such a retrieval paradigm, an end user searches for unlabeled videos by ad-hoc queries described in natural language text with no visual example provided. Given videos as sequences of frames and queries as sequences of words, an effective sequence-to-sequence cross-modal matching is required. The majority of existing methods are concept based, extracting relevant concepts from queries and videos and accordingly establishing associations between the two modalities. In contrast, this paper takes a concept-free approach, proposing a dual deep encoding network that encodes videos and queries into powerful dense representations of their own. Dual encoding is conceptually simple, practically effective and end-to-end. As experiments on three benchmarks, i.e. MSR-VTT, TRECVID 2016 and 2017 Ad-hoc Video Search show, the proposed solution establishes a new state-of-the-art for zero-example video retrieval.

📄 PDF Abstract BibTeX arXiv:1809.06181

Code (1)

danieljf24/dual_encoding 공식 구현 pytorch

Tasks

Ad-hoc video searchRetrievalVideo Retrieval

Similar Papers 제목 키워드 기반

Dual Encoding for Video Retrieval by Text

2020-09-10 · Jianfeng Dong, Xirong Li, Chaoxi Xu, Xun Yang 외

This paper attacks the challenging problem of video retrieval by text. In such a retrieval paradigm, an end user searches for unlabeled videos by ad-hoc queries described exclusively in the form of a natural-language sen…

Ad-hoc video searchRetrievalSentenceVideo Retrieval

Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval

2025-03-24 · CVPR 2025 1 · Arun Reddy, Alexander Martin, Eugene Yang, Andrew Yates 외

In this work, we tackle the problem of text-to-video retrieval (T2VR). Inspired by the success of late interaction techniques in text-document, text-image, and text-video retrieval, our approach, Video-ColBERT, introduce…

RetrievalText to Video RetrievalVideo Retrieval

Everything at Once -- Multi-modal Fusion Transformer for Video Retrieval

2021-12-08 · Nina Shvetsova, Brian Chen, Andrew Rouditchenko, Samuel Thomas 외

Multi-modal learning from video data has seen increased attention recently as it allows to train semantically meaningful embeddings without human annotation enabling tasks like zero-shot retrieval and classification. In …

Action LocalizationRetrievalVideo RetrievalZero-Shot Video Retrieval

VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

2021-09-28 · EMNLP 2021 11 · Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko 외

We present VideoCLIP, a contrastive approach to pre-train a unified model for zero-shot video and text understanding, without using any labels on downstream tasks. VideoCLIP trains a transformer for video and text by con…

Action LocalizationAction SegmentationLong Video Retrieval (Background Removed)Retrieval+4

Everything at Once - Multi-Modal Fusion Transformer for Video Retrieval

2022-01-01 · CVPR 2022 1 · Nina Shvetsova, Brian Chen, Andrew Rouditchenko, Samuel Thomas 외

Multi-modal learning from video data has seen increased attention recently as it allows training of semantically meaningful embeddings without human annotation, enabling tasks like zero-shot retrieval and action loca…

Action LocalizationRetrievalVideo RetrievalZero-Shot Video Retrieval