paper-with-me

Papers

Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse

2025-06-17 · Jinwoo Hwang, Daeun Kim, Sangyeop Lee, Yoonsung Kim, Guseul Heo, Hojoon Kim, Yunseok Jeong, Tadiwos Meaza, Eunhyeok Park, Jeongseob Ahn, Jongse Park

Recently, Video-Language Models (VideoLMs) have demonstrated remarkable capabilities, offering significant potential for flexible and powerful video query systems. These models typically rely on Vision Transformers (ViTs), which process video frames individually to extract visual embeddings. However, generating embeddings for large-scale videos requires ViT inferencing across numerous frames, posing a major hurdle to real-world deployment and necessitating solutions for integration into scalable video data management systems. This paper introduces D\'ej\a Vu, a video-language query engine that accelerates ViT-based VideoLMs by reusing computations across consecutive frames. At its core is ReuseViT, a modified ViT model specifically designed for VideoLM tasks, which learns to detect inter-frame reuse opportunities, striking an effective balance between accuracy and reuse. Although ReuseViT significantly reduces computation, these savings do not directly translate into performance gains on GPUs. To overcome this, D\'ej\a Vu integrates memory-compute joint compaction techniques that convert the FLOP savings into tangible performance gains. Evaluations on three VideoLM tasks show that D\'ej\`a Vu accelerates embedding generation by up to a 2.64x within a 2% error bound, dramatically enhancing the practicality of VideoLMs for large-scale video analytics.

📄 PDF Abstract BibTeX arXiv:2506.14107

Code (1)

casys-kaist/dejavu 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Towards A Time Based Video Search Engine for Al Quran Interpretation

2017-01-25 · Eljazzar Maged M., Hassan Afnan, AlSharkawy Amira A.

The number of Internet Muslim-users is remarkably increasing from all over the world countries. There are a lot of structured, and well-documented text resources for the Quran interpretation, Tafsir, over the Internet wi…

Beyond Single-Modal Analytics: A Framework for Integrating Heterogeneous LLM-Based Query Systems for Multi-Modal Data

2026-02-02 · Ruyu Li, Tinghui Zhang, Haodi Ma, Daisy Zhe Wang 외 arxiv

With the increasing use of multi-modal data, semantic query has become more and more demanded in data management systems, which is an important way to access and analyze multi-modal data. As unstructured data, most infor…

VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement

2026-07-01 · Seohyun Lee, Seoung Choi, Dohwan Ko, Jongha Kim 외 hf

As video corpora continue to expand in both scale and task complexity, there is increasing demand for approaches that retrieve relevant videos from large-scale corpora (inter-video reasoning) and subsequently perform fin…

Moment RetrievalVideo Retrieval

IVSS Integration of Color Feature Extraction Techniques for Intelligent Video Search Systems

2013-12-24 · Avinash N Bhute, B. B. Meshram

As large amount of visual Information is available on web in form of images, graphics, animations and videos, so it is important in internet era to have an effective video search system. As there are number of video sear…

Tree-based Text-Vision BERT for Video Search in Baidu Video Advertising

2022-09-19 · Tan Yu, Jie Liu, Yi Yang, Yi Li 외

The advancement of the communication technology and the popularity of the smart phones foster the booming of video ads. Baidu, as one of the leading search engine companies in the world, receives billions of search queri…

Image RetrievalRetrievalVideo Retrieval