paper-with-me

홈 › Papers

HORNet: Task-Guided Frame Selection for Video Question Answering with Vision-Language Models

2026-03-19 · Xiangyu Bai, Bishoy Galoaa, Sarah Ostadabbas arxiv

Video question answering (VQA) with vision-language models (VLMs) depends critically on which frames are selected from the input video, yet most systems rely on uniform or heuristic sampling that cannot be optimized for downstream answering quality. We introduce \textbf{HORNet}, a lightweight frame selection policy trained with Group Relative Policy Optimization (GRPO) to learn which frames a frozen VLM needs to answer questions correctly. With fewer than 1M trainable parameters, HORNet reduces input frames by up to 99\% and VLM processing time by up to 93\%, while improving answer quality on short-form benchmarks (+1.7\% F1 on MSVD-QA) and achieving strong performance on temporal reasoning tasks (+7.3 points over uniform sampling on NExT-QA). We formalize this as Select Any Frames (SAF), a task that decouples visual input curation from VLM reasoning, and show that GRPO-trained selection generalizes better out-of-distribution than supervised and PPO alternatives. HORNet's policy further transfers across VLM answerers without retraining, yielding an additional 8.5\% relative gain when paired with a stronger model. Evaluated across six benchmarks spanning 341,877 QA pairs and 114.2 hours of video, our results demonstrate that optimizing \emph{what} a VLM sees is a practical and complementary alternative to optimizing what it generates while improving efficiency. Code is available at https://github.com/ostadabbas/HORNet.

📄 PDF Abstract BibTeX arXiv:2603.18850

Code (0)

등록된 구현이 없습니다.

Tasks

Video Question Answering

Similar Papers 제목 키워드 기반

HorNet: Efficient High-Order Spatial Interactions with Recursive Gated Convolutions

2022-07-28 · Yongming Rao, Wenliang Zhao, Yansong Tang, Jie zhou 외

Recent progress in vision Transformers exhibits great success in various tasks driven by the new spatial modeling mechanism based on dot-product self-attention. In this paper, we show that the key ingredients behind the …

Image ClassificationObject DetectionSemantic SegmentationVocal Bursts Intensity Prediction

Localizing Semantic Patches for Accelerating Image Classification

2022-06-07 · Chuanguang Yang, Zhulin An, Yongjun Xu

Existing works often focus on reducing the architecture redundancy for accelerating image classification but ignore the spatial redundancy of the input image. This paper proposes an efficient image classification pipelin…

ClassificationGeneral Classificationimage-classificationImage Classification

HorNets: Learning from Discrete and Continuous Signals with Routing Neural Networks

2025-01-24 · Boshko Koloski, Nada Lavrač, Blaž Škrlj

Construction of neural network architectures suitable for learning from both continuous and discrete tabular data is a challenging research endeavor. Contemporary high-dimensional tabular data sets are often characterize…

Priority prediction of Asian Hornet sighting report using machine learning methods

2021-06-28 · Yixin Liu, Jiaxin Guo, Jieyang Dong, Luoqian Jiang 외

As infamous invaders to the North American ecosystem, the Asian giant hornet (Vespa mandarinia) is devastating not only to native bee colonies, but also to local apiculture. One of the most effective way to combat the ha…

BIG-bench Machine Learning

HorNet: A Hierarchical Offshoot Recurrent Network for Improving Person Re-ID via Image Captioning

2019-08-14 · Shi-Yang Yan, Jun Xu, Yuai Liu, Lin Xu

Person re-identification (re-ID) aims to recognize a person-of-interest across different cameras with notable appearance variance. Existing research works focused on the capability and robustness of visual representation…

Generative Adversarial NetworkImage CaptioningPerson Re-Identification