Papers Video to Text Retrieval
“Video to Text Retrieval” 태그가 달린 논문 13편 · 필터 해제
SPECTRUM: Semantic Processing and Emotion-informed video-Captioning Through Retrieval and Understanding Modalities
Capturing a video's meaning and critical concepts by analyzing the subtle details is a fundamental yet challenging task in video captioning. Identifying the dominant emotional tone in a video significantly enhances the p…
AttributeDescriptiveRetrievalText Retrieval+2COM Kitchens: An Unedited Overhead-view Video Dataset as a Vision-Language Benchmark
Procedural video understanding is gaining attention in the vision and language community. Deep learning-based video analysis requires extensive data. Consequently, existing works often use web videos as training resource…
Dense Video CaptioningDiversityRetrievalText Retrieval+3SignCLIP: Connecting Text and Sign Language by Contrastive Learning
We present SignCLIP, which re-purposes CLIP (Contrastive Language-Image Pretraining) to project spoken language text and sign language videos, two classes of natural languages of distinct modalities, into the same space.…
Contrastive LearningRetrievalSign Language RecognitionText Retrieval+1Sakuga-42M Dataset: Scaling Up Cartoon Research
Hand-drawn cartoon animation employs sketches and flat-color segments to create the illusion of motion. While recent advancements like CLIP, SVD, and Sora show impressive results in understanding and generating natural v…
MambaText to Video RetrievalVideo to Text RetrievalPrototype-based Aleatoric Uncertainty Quantification for Cross-modal Retrieval
Cross-modal Retrieval methods build similarity relations between vision and language modalities by jointly learning a common representation space. However, the predictions are often unreliable due to the Aleatoric uncert…
Cross-Modal RetrievalImage-text matchingImage-to-Text RetrievalRetrieval+6MSVD-Indonesian: A Benchmark for Multimodal Video-Text Tasks in Indonesian
Multimodal learning on video and text data has been receiving growing attention from many researchers in various research tasks, including text-to-video retrieval, video-to-text retrieval, and video captioning. Although …
Cross-Lingual TransferRetrievalText RetrievalText to Video Retrieval+5i-Code Studio: A Configurable and Composable Framework for Integrative AI
Artificial General Intelligence (AGI) requires comprehensive understanding and generation capabilities for a variety of tasks spanning different modalities and functionalities. Integrative AI is one important direction t…
Question AnsweringRetrievalSpeech-to-Speech TranslationText Retrieval+2VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners
We explore an efficient approach to establish a foundational video-text model. We present VideoCoCa that maximally reuses a pretrained image-text contrastive captioner (CoCa) model and adapt it to video-text tasks with m…
Question AnsweringRetrievalText to Video RetrievalVideo Captioning+7MILES: Visual BERT Pre-training with Injected Language Semantics for Video-text Retrieval
Dominant pre-training work for video-text retrieval mainly adopt the "dual-encoder" architectures to enable efficient retrieval, where two separate encoders are used to contrast global video and text representations, but…
Action RecognitionRetrievalText RetrievalText to Video Retrieval+5Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language
Large pretrained (e.g., "foundation") models exhibit distinct capabilities depending on the domain of data they are trained on. While these domains are generic, they may only barely overlap. For example, visual-language …
DiversityImage CaptioningMultimodal ReasoningRetrieval+4Bridging Video-text Retrieval with Multiple Choice Questions
Pre-training a model to learn transferable video-text representation for retrieval has attracted a lot of attention in recent years. Previous dominant works mainly adopt two separate encoders for efficient retrieval, but…
Action RecognitionLinear evaluationMultiple-choiceRetrieval+8CLIP2Video: Mastering Video-Text Retrieval via Image CLIP
We present CLIP2Video network to transfer the image-language pre-training model to video-text retrieval in an end-to-end manner. Leading approaches in the domain of video-and-language learning try to distill the spatio-t…
Language ModelingLanguage ModellingRetrievalText Retrieval+3Learning a Text-Video Embedding from Incomplete and Heterogeneous Data
Joint understanding of video and language is an active research area with many applications. Prior work in this domain typically relies on learning text-video embeddings. One difficulty with this approach, however, is th…
RetrievalText RetrievalVideo RetrievalVideo to Text Retrieval