Enhancing Subsequent Video Retrieval via Vision-Language Models (VLMs)
The rapid growth of video content demands efficient and precise retrieval systems. While vision-language models (VLMs) excel in representation learning, they often struggle with adaptive, time-sensitive video retrieval. This paper introduces a novel framework that combines vector similarity search with graph-based data structures. By leveraging VLM embeddings for initial retrieval and modeling contextual relationships among video segments, our approach enables adaptive query refinement and improves retrieval accuracy. Experiments demonstrate its precision, scalability, and robustness, offering an effective solution for interactive video retrieval in dynamic environments.
Code (1)
Tasks
Representation LearningRetrievalVideo RetrievalSimilar Papers 제목 키워드 기반
V-Agent: An Interactive Video Search System Using Vision-Language Models
We introduce V-Agent, a novel multi-agent platform designed for advanced video search and interactive user-system conversations. By fine-tuning a vision-language model (VLM) with a small video preference dataset and enha…
Speech RecognitionVideo RetrievalText RetrievalTowards Effective Long-Video Event Prediction via Multi-Level Event Semantics Mining
Accurately predicting future events is fundamental to content understanding and decision-making across various domains. While prior research has primarily focused on text or short-video scenarios, long-video event predic…
Video Question AnsweringLearning to Retrieve Videos by Asking Questions
The majority of traditional text-to-video retrieval systems operate in static environments, i.e., there is no interaction between the user and the agent beyond the initial textual query provided by the user. This can be …
AI AgentRetrievalText to Video RetrievalVideo RetrievalE-VRAG: Enhancing Long Video Understanding with Resource-Efficient Retrieval Augmented Generation
Vision-Language Models (VLMs) have enabled substantial progress in video understanding by leveraging cross-modal reasoning capabilities. However, their effectiveness is limited by the restricted context window and the hi…
Question AnsweringVideo sentence grounding with temporally global textual knowledge
Temporal sentence grounding involves the retrieval of a video moment with a natural language query. Many existing works directly incorporate the given video and temporally localized query for temporal grounding, overlook…
Contrastive LearningRetrievalSentenceTemporal Sentence Grounding