paper-with-me

홈 › Papers

Enhancing Subsequent Video Retrieval via Vision-Language Models (VLMs)

2025-03-21 · Yicheng Duan, Xi Huang, Duo Chen

The rapid growth of video content demands efficient and precise retrieval systems. While vision-language models (VLMs) excel in representation learning, they often struggle with adaptive, time-sensitive video retrieval. This paper introduces a novel framework that combines vector similarity search with graph-based data structures. By leveraging VLM embeddings for initial retrieval and modeling contextual relationships among video segments, our approach enables adaptive query refinement and improves retrieval accuracy. Experiments demonstrate its precision, scalability, and robustness, offering an effective solution for interactive video retrieval in dynamic environments.

📄 PDF Abstract BibTeX arXiv:2503.17415

Code (1)

YichengDuan/svrllm 공식 구현

Tasks

Representation LearningRetrievalVideo Retrieval

Similar Papers 제목 키워드 기반

V-Agent: An Interactive Video Search System Using Vision-Language Models

2025-11-04 · SunYoung Park, Jong-Hyeon Lee, Youngjune Kim, Daegyu Sung 외 arxiv

We introduce V-Agent, a novel multi-agent platform designed for advanced video search and interactive user-system conversations. By fine-tuning a vision-language model (VLM) with a small video preference dataset and enha…

Speech RecognitionVideo RetrievalText Retrieval

Towards Effective Long-Video Event Prediction via Multi-Level Event Semantics Mining

2026-05-29 · Bo Peng, YuanJie Lyu, PengGang Qin, Tong Xu arxiv

Accurately predicting future events is fundamental to content understanding and decision-making across various domains. While prior research has primarily focused on text or short-video scenarios, long-video event predic…

Video Question Answering

Learning to Retrieve Videos by Asking Questions

2022-05-11 · Avinash Madasu, Junier Oliva, Gedas Bertasius

The majority of traditional text-to-video retrieval systems operate in static environments, i.e., there is no interaction between the user and the agent beyond the initial textual query provided by the user. This can be …

AI AgentRetrievalText to Video RetrievalVideo Retrieval

E-VRAG: Enhancing Long Video Understanding with Resource-Efficient Retrieval Augmented Generation

2025-08-03 · Zeyu Xu, Junkang Zhang, Qiang Wang, Yi Liu arxiv

Vision-Language Models (VLMs) have enabled substantial progress in video understanding by leveraging cross-modal reasoning capabilities. However, their effectiveness is limited by the restricted context window and the hi…

Question Answering

Video sentence grounding with temporally global textual knowledge

2024-04-21 · Cai Chen, Runzhong Zhang, Jianjun Gao, Kejun Wu 외

Temporal sentence grounding involves the retrieval of a video moment with a natural language query. Many existing works directly incorporate the given video and temporally localized query for temporal grounding, overlook…

Contrastive LearningRetrievalSentenceTemporal Sentence Grounding