HaVTR: Improving Video-Text Retrieval Through Augmentation Using Large Foundation Models
While recent progress in video-text retrieval has been driven by the exploration of powerful model architectures and training strategies, the representation learning ability of video-text retrieval models is still limited due to low-quality and scarce training data annotations. To address this issue, we present a novel video-text learning paradigm, HaVTR, which augments video and text data to learn more generalized features. Specifically, we first adopt a simple augmentation method, which generates self-similar data by randomly duplicating or dropping subwords and frames. In addition, inspired by the recent advancement in visual and language generative models, we propose a more powerful augmentation method through textual paraphrasing and video stylization using large language models (LLMs) and visual generative models (VGMs). Further, to bring richer information into video and text, we propose a hallucination-based augmentation method, where we use LLMs and VGMs to generate and add new relevant information to the original data. Benefiting from the enriched data, extensive experiments on several video-text retrieval benchmarks demonstrate the superiority of HaVTR over existing methods.
Code (0)
등록된 구현이 없습니다.
Tasks
HallucinationRepresentation LearningRetrievalText RetrievalVideo-Text RetrievalSimilar Papers 제목 키워드 기반
Multi-Modal Retrieval Augmentation for Open-Ended and Knowledge-Intensive Video Question Answering
While current video question answering systems perform well on some tasks requiring only direct visual understanding, they struggle with questions demanding knowledge beyond what is immediately observable in the video co…
Multiple-choiceQuestion AnsweringRetrievalRetrieval-augmented Generation+2EA-VTR: Event-Aware Video-Text Retrieval
Understanding the content of events occurring in the video and their inherent temporal logic is crucial for video-text retrieval. However, web-crawled pre-training datasets often lack sufficient event information, and th…
Action RecognitionContrastive Learningcross-modal alignmentMoment Retrieval+6A Feature-space Multimodal Data Augmentation Technique for Text-video Retrieval
Every hour, huge amounts of visual contents are posted on social media and user-generated content platforms. To find relevant videos by means of a natural language query, text-video retrieval methods have received increa…
Data AugmentationRetrievalVideo RetrievalT2VIndexer: A Generative Video Indexer for Efficient Text-Video Retrieval
Current text-video retrieval methods mainly rely on cross-modal matching between queries and videos to calculate their similarity scores, which are then sorted to obtain retrieval results. This method considers the match…
RetrievalVideo RetrievalExpertized Caption Auto-Enhancement for Video-Text Retrieval
Video-text retrieval has been stuck in the information mismatch caused by personalized and inadequate textual descriptions of videos. The substantial information gap between the two modalities hinders an effective cross-…
Caption GenerationRetrievalText RetrievalVideo-Text Retrieval