paper-with-me

홈 › Papers

Generative Frame Sampler for Long Video Understanding

2025-03-12 · Linli Yao, HaoNing Wu, Kun Ouyang, Yuanxing Zhang, Caiming Xiong, Bei Chen, Xu sun, Junnan Li

Despite recent advances in Video Large Language Models (VideoLLMs), effectively understanding long-form videos remains a significant challenge. Perceiving lengthy videos containing thousands of frames poses substantial computational burden. To mitigate this issue, this paper introduces Generative Frame Sampler (GenS), a plug-and-play module integrated with VideoLLMs to facilitate efficient lengthy video perception. Built upon a lightweight VideoLLM, GenS leverages its inherent vision-language capabilities to identify question-relevant frames. To facilitate effective retrieval, we construct GenS-Video-150K, a large-scale video instruction dataset with dense frame relevance annotations. Extensive experiments demonstrate that GenS consistently boosts the performance of various VideoLLMs, including open-source models (Qwen2-VL-7B, Aria-25B, VILA-40B, LLaVA-Video-7B/72B) and proprietary assistants (GPT-4o, Gemini). When equipped with GenS, open-source VideoLLMs achieve impressive state-of-the-art results on long-form video benchmarks: LLaVA-Video-72B reaches 66.8 (+4.3) on LongVideoBench and 77.0 (+2.7) on MLVU, while Aria obtains 39.2 on HourVideo surpassing the Gemini-1.5-pro by 1.9 points. We will release all datasets and models at https://generative-sampler.github.io.

📄 PDF Abstract BibTeX arXiv:2503.09146

Code (0)

등록된 구현이 없습니다.

Tasks

Video Understanding

Methods 이 논문이 사용한 방법론

ARiA 설명 없음

Similar Papers 제목 키워드 기반

MSJoE: Jointly Evolving MLLM and Sampler for Efficient Long-Form Video Understanding

2026-02-26 · Wenhui Tan, Xiaoyi Yu, Jiaze Li, Yijing Chen 외 arxiv

Efficiently understanding long-form videos remains a fundamental challenge for multimodal large language models (MLLMs). In this paper, we present MLLM-Sampler Joint Evolution (MSJoE), a novel framework that jointly evol…

Reinforcement LearningAnswer Generation

LongVU-TTT: Causal Test-Time Training for Visual Resampling in Long Video Understanding

2026-08-26 · Mahmoud Ahmed, Sameh Abdulah, Olatunji Ruwase, Sam Ade Jacobs 외 arxiv

Long-video MLLMs must model temporal change before a limited visual-token budget removes most frame evidence. We introduce LongVU-TTT, which inserts a convolutional Test-Time Training (TTT) resampler with causal fast-wei…

MuLTI: Efficient Video-and-Language Understanding with Text-Guided MultiWay-Sampler and Multiple Choice Modeling

2023-03-10 · Jiaqi Xu, Bo Liu, Yunkuo Chen, Mengli Cheng 외

Video-and-language understanding has a variety of applications in the industry, such as video question answering, text-video retrieval, and multi-label classification. Existing video-and-language understanding methods ge…

Multi-Label ClassificationMUlTI-LABEL-ClASSIFICATIONMultiple-choiceQuestion Answering+7

Shot-Aware Frame Sampling for Video Understanding

2026-03-18 · Mengyu Zhao, Di Fu, Yongyu Xie, Jiaxing Zhang 외 arxiv

Video frame sampling is essential for efficient long-video understanding with Vision-Language Models (VLMs), since dense inputs are costly and often exceed context limits. Yet when only a small number of frames can be re…

Text-Conditioned Resampler For Long Form Video Understanding

2023-12-19 · Bruno Korbar, Yongqin Xian, Alessio Tonioni, Andrew Zisserman 외

In this paper we present a text-conditioned video resampler (TCR) module that uses a pre-trained and frozen visual encoder and large language model (LLM) to process long video sequences for a task. TCR localises relevant…

EgoSchemaFormLanguage ModelingLanguage Modelling+3