paper-with-me

Papers Video to Text Retrieval

“Video to Text Retrieval” 태그가 달린 논문 13편 · 필터 해제

SPECTRUM: Semantic Processing and Emotion-informed video-Captioning Through Retrieval and Understanding Modalities

2024-11-04 · Ehsan Faghihi, Mohammedreza Zarenejad, Ali-Asghar Beheshti Shirazi

Capturing a video's meaning and critical concepts by analyzing the subtle details is a fundamental yet challenging task in video captioning. Identifying the dominant emotional tone in a video significantly enhances the p…

AttributeDescriptiveRetrievalText Retrieval+2

COM Kitchens: An Unedited Overhead-view Video Dataset as a Vision-Language Benchmark

2024-08-05 · Koki Maeda, Tosho Hirasawa, Atsushi Hashimoto, Jun Harashima 외

Procedural video understanding is gaining attention in the vision and language community. Deep learning-based video analysis requires extensive data. Consequently, existing works often use web videos as training resource…

Dense Video CaptioningDiversityRetrievalText Retrieval+3

SignCLIP: Connecting Text and Sign Language by Contrastive Learning

2024-07-01 · Zifan Jiang, Gerard Sant, Amit Moryossef, Mathias Müller 외

We present SignCLIP, which re-purposes CLIP (Contrastive Language-Image Pretraining) to project spoken language text and sign language videos, two classes of natural languages of distinct modalities, into the same space.…

Contrastive LearningRetrievalSign Language RecognitionText Retrieval+1

Sakuga-42M Dataset: Scaling Up Cartoon Research

2024-05-13 · Zhenglin Pan, Yu Zhu, Yuxuan Mu

Hand-drawn cartoon animation employs sketches and flat-color segments to create the illusion of motion. While recent advancements like CLIP, SVD, and Sora show impressive results in understanding and generating natural v…

MambaText to Video RetrievalVideo to Text Retrieval

Prototype-based Aleatoric Uncertainty Quantification for Cross-modal Retrieval

2023-09-29 · NeurIPS 2023 11 · Hao Li, Jingkuan Song, Lianli Gao, Xiaosu Zhu 외

Cross-modal Retrieval methods build similarity relations between vision and language modalities by jointly learning a common representation space. However, the predictions are often unreliable due to the Aleatoric uncert…

Cross-Modal RetrievalImage-text matchingImage-to-Text RetrievalRetrieval+6

MSVD-Indonesian: A Benchmark for Multimodal Video-Text Tasks in Indonesian

2023-06-20 · Willy Fitra Hendria

Multimodal learning on video and text data has been receiving growing attention from many researchers in various research tasks, including text-to-video retrieval, video-to-text retrieval, and video captioning. Although …

Cross-Lingual TransferRetrievalText RetrievalText to Video Retrieval+5

i-Code Studio: A Configurable and Composable Framework for Integrative AI

2023-05-23 · Yuwei Fang, Mahmoud Khademi, Chenguang Zhu, ZiYi Yang 외

Artificial General Intelligence (AGI) requires comprehensive understanding and generation capabilities for a variety of tasks spanning different modalities and functionalities. Integrative AI is one important direction t…

Question AnsweringRetrievalSpeech-to-Speech TranslationText Retrieval+2

VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners

2022-12-09 · Shen Yan, Tao Zhu, ZiRui Wang, Yuan Cao 외

We explore an efficient approach to establish a foundational video-text model. We present VideoCoCa that maximally reuses a pretrained image-text contrastive captioner (CoCa) model and adapt it to video-text tasks with m…

Question AnsweringRetrievalText to Video RetrievalVideo Captioning+7

MILES: Visual BERT Pre-training with Injected Language Semantics for Video-text Retrieval

2022-04-26 · Yuying Ge, Yixiao Ge, Xihui Liu, Alex Jinpeng Wang 외

Dominant pre-training work for video-text retrieval mainly adopt the "dual-encoder" architectures to enable efficient retrieval, where two separate encoders are used to contrast global video and text representations, but…

Action RecognitionRetrievalText RetrievalText to Video Retrieval+5

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language

2022-04-01 · Andy Zeng, Maria Attarian, Brian Ichter, Krzysztof Choromanski 외

Large pretrained (e.g., "foundation") models exhibit distinct capabilities depending on the domain of data they are trained on. While these domains are generic, they may only barely overlap. For example, visual-language …

DiversityImage CaptioningMultimodal ReasoningRetrieval+4

Bridging Video-text Retrieval with Multiple Choice Questions

2022-01-13 · CVPR 2022 1 · Yuying Ge, Yixiao Ge, Xihui Liu, Dian Li 외

Pre-training a model to learn transferable video-text representation for retrieval has attracted a lot of attention in recent years. Previous dominant works mainly adopt two separate encoders for efficient retrieval, but…

Action RecognitionLinear evaluationMultiple-choiceRetrieval+8

CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

2021-06-21 · Han Fang, Pengfei Xiong, Luhui Xu, Yu Chen

We present CLIP2Video network to transfer the image-language pre-training model to video-text retrieval in an end-to-end manner. Leading approaches in the domain of video-and-language learning try to distill the spatio-t…

Language ModelingLanguage ModellingRetrievalText Retrieval+3

Learning a Text-Video Embedding from Incomplete and Heterogeneous Data

2018-04-07 · Antoine Miech, Ivan Laptev, Josef Sivic

Joint understanding of video and language is an active research area with many applications. Prior work in this domain typically relies on learning text-video embeddings. One difficulty with this approach, however, is th…

RetrievalText RetrievalVideo RetrievalVideo to Text Retrieval
1–13 / 13