paper-with-me

홈 › Papers

Efficient Video Sampling: Pruning Temporally Redundant Tokens for Faster VLM Inference

2025-10-16 · Natan Bagrov, Eugene Khvedchenia, Borys Tymchenko, Shay Aharon, Lior Kadoch, Tomer Keren, Ofri Masad, Yonatan Geifman, Ran Zilberstein, Tuomas Rintamaki, Matthieu Le, Andrew Tao arxiv

Vision-language models (VLMs) have recently expanded from static image understanding to video reasoning, but their scalability is fundamentally limited by the quadratic cost of processing dense frame sequences. Long videos often exceed the token budget of modern language models, leading to severe context limitations and latency issues. We introduce Efficient Video Sampling (EVS), a simple, plug-and-play method for reducing token redundancy in videos by identifying and pruning temporally static patches -- spatial regions that remain unchanged across consecutive frames. EVS preserves positional identity, requires no architectural changes or retraining. We show that EVS substantially reduces token count while maintaining semantic fidelity, enabling faster inference and longer input sequences. Applied at inference time, EVS reduces large language model (LLM) time-to-first-token (TTFT) by up to 4x with minimal accuracy loss. When combined with an uptraining phase using stochastic pruning rates, EVS yields models that are robust to varying compression levels and retain full performance under aggressive pruning. Extensive experiments demonstrate that EVS consistently improves efficiency-accuracy trade-offs, unlocking scalable video-language understanding without sacrificing quality.

📄 PDF Abstract BibTeX arXiv:2510.14624

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent

2025-07-21 · Jiaao Li, Kaiyuan Li, Chen Gao, Yong Li 외

Egomotion videos are first-person recordings where the view changes continuously due to the agent's movement. As they serve as the primary visual input for embodied AI agents, making egomotion video reasoning more effici…

Multimodal Reasoning

EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs

2026-05-11 · Jiameng Li, Minye Wu, Jiezhang Cao, Aleksei Tiulpin 외 arxiv

Long-form video understanding remains challenging for Video Large Language Models (VideoLLMs), as the dense frame sampling introduces massive visual tokens while sparse sampling risks missing critical temporal evidence a…

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models

2026-08-04 · Mengjie Zhang, Qihui Zhu, Tao Zhang, Shuangwu Chen 외 arxiv

Video large language models (VideoLLMs) achieve strong video understanding performance, but their inference remains expensive due to the large number of redundant spatio-temporal visual tokens in long videos. Existing to…

FastVID: Dynamic Density Pruning for Fast Video Large Language Models

2025-03-14 · Leqi Shen, Guoqiang Gong, Tao He, Yifeng Zhang 외

Video Large Language Models have shown impressive capabilities in video comprehension, yet their practical deployment is hindered by substantial inference costs caused by redundant video tokens. Existing pruning techniqu…

CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models

2025-08-24 · Zicong Tang, Ziyang Ma, Suqing Wang, Zuchao Li 외 arxiv

Large Vision-Language Models (LVLMs) process multimodal inputs consisting of text tokens and vision tokens extracted from images or videos. Due to the rich visual information, a single image can generate thousands of vis…