paper-with-me

홈 › Papers

HaVTR: Improving Video-Text Retrieval Through Augmentation Using Large Foundation Models

2024-04-07 · Yimu Wang, Shuai Yuan, Xiangru Jian, Wei Pang, Mushi Wang, Ning Yu

While recent progress in video-text retrieval has been driven by the exploration of powerful model architectures and training strategies, the representation learning ability of video-text retrieval models is still limited due to low-quality and scarce training data annotations. To address this issue, we present a novel video-text learning paradigm, HaVTR, which augments video and text data to learn more generalized features. Specifically, we first adopt a simple augmentation method, which generates self-similar data by randomly duplicating or dropping subwords and frames. In addition, inspired by the recent advancement in visual and language generative models, we propose a more powerful augmentation method through textual paraphrasing and video stylization using large language models (LLMs) and visual generative models (VGMs). Further, to bring richer information into video and text, we propose a hallucination-based augmentation method, where we use LLMs and VGMs to generate and add new relevant information to the original data. Benefiting from the enriched data, extensive experiments on several video-text retrieval benchmarks demonstrate the superiority of HaVTR over existing methods.

📄 PDF Abstract BibTeX arXiv:2404.05083

Code (0)

등록된 구현이 없습니다.

Tasks

HallucinationRepresentation LearningRetrievalText RetrievalVideo-Text Retrieval

Similar Papers 제목 키워드 기반

Multi-Modal Retrieval Augmentation for Open-Ended and Knowledge-Intensive Video Question Answering

2025-02-17 · Md Zarif Ul Alam, Hamed Zamani

While current video question answering systems perform well on some tasks requiring only direct visual understanding, they struggle with questions demanding knowledge beyond what is immediately observable in the video co…

Multiple-choiceQuestion AnsweringRetrievalRetrieval-augmented Generation+2

EA-VTR: Event-Aware Video-Text Retrieval

2024-07-10 · Zongyang Ma, Ziqi Zhang, Yuxin Chen, Zhongang Qi 외

Understanding the content of events occurring in the video and their inherent temporal logic is crucial for video-text retrieval. However, web-crawled pre-training datasets often lack sufficient event information, and th…

Action RecognitionContrastive Learningcross-modal alignmentMoment Retrieval+6

A Feature-space Multimodal Data Augmentation Technique for Text-video Retrieval

2022-08-03 · Alex Falcon, Giuseppe Serra, Oswald Lanz

Every hour, huge amounts of visual contents are posted on social media and user-generated content platforms. To find relevant videos by means of a natural language query, text-video retrieval methods have received increa…

Data AugmentationRetrievalVideo Retrieval

T2VIndexer: A Generative Video Indexer for Efficient Text-Video Retrieval

2024-08-21 · Yili Li, Jing Yu, Keke Gai, Bang Liu 외

Current text-video retrieval methods mainly rely on cross-modal matching between queries and videos to calculate their similarity scores, which are then sorted to obtain retrieval results. This method considers the match…

RetrievalVideo Retrieval

Expertized Caption Auto-Enhancement for Video-Text Retrieval

2025-02-05 · Baoyao Yang, Junxiang Chen, Wanyun Li, Wenbin Yao 외

Video-text retrieval has been stuck in the information mismatch caused by personalized and inadequate textual descriptions of videos. The substantial information gap between the two modalities hinders an effective cross-…

Caption GenerationRetrievalText RetrievalVideo-Text Retrieval