paper-with-me

Papers

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models

2026-08-04 · Mengjie Zhang, Qihui Zhu, Tao Zhang, Shuangwu Chen, Huihuang Qin, Yu Guo, Shenghao Ye, Zijian Wen, Yunpeng Hou, Dong Jin, Xiaobin Tan, Huasen He, Jian Yang arxiv

Video large language models (VideoLLMs) achieve strong video understanding performance, but their inference remains expensive due to the large number of redundant spatio-temporal visual tokens in long videos. Existing token pruning methods alleviate this cost by reducing redundant tokens, yet most of them rely on segment-level local pruning, where videos are partitioned into isolated segments and tokens are selected independently within each segment. Such designs may under-preserve short but semantically dense segments and discard tokens that appear non-salient locally but remain critical from a global perspective. To address this issue, we propose GSTEP (Global Spatio-Temporal Density Pruning), a plug-and-play pruning framework that models video as a continuous spatio-temporal information flow. GSTEP constructs a token-level spatio-temporal density by combining a continuous temporal density, obtained from a smoothed centered frame-level change signal, with intra-frame spatial density, and then performs global token sampling by jointly balancing information density and coverage. Extensive experiments on multiple VideoLLMs and public benchmarks demonstrate that GSTEP consistently achieves strong accuracy-efficiency trade-offs and generalizes well across model architectures and evaluation settings. On LLaVA-OneVision-7B, GSTEP prunes 75% of visual tokens, preserves up to 100.2% of the original average performance across benchmarks, and achieves a 1.17 end-to-end speedup.

📄 PDF Abstract BibTeX arXiv:2608.03083

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

2023-11-09 · Yichen Gong, Delong Ran, JinYuan Liu, Conglei Wang 외

Large Vision-Language Models (LVLMs) signify a groundbreaking paradigm shift within the Artificial Intelligence (AI) community, extending beyond the capabilities of Large Language Models (LLMs) by assimilating additional…

Optical Character Recognition (OCR)Safety Alignment

Physics-guided spatiotemporal neural models for fuel density prediction

2026-07-08 · Tolga Caglar, Jaynil Jaiswal, Saqib Azim, Yudhir Gala 외 arxiv

This paper presents a physics-guided machine learning (PGML) framework for fuel density prediction, integrating physics constraints and domain knowledge into deep learning models to enhance model accuracy and stability. …

Global Stress Generation and Spatiotemporal Super-Resolution Physics-Informed Operator under Dynamic Loading for Two-Phase Random Materials

2025-04-26 · Tengfei Xing, Xiaodan Ren, Jie Li

Material stress analysis is a critical aspect of material design and performance optimization. Under dynamic loading, the global stress evolution in materials exhibits complex spatiotemporal characteristics, especially i…

STSSuper-Resolution

Multi-Spatio-temporal Fusion Graph Recurrent Network for Traffic forecasting

2022-05-03 · Wei Zhao, Shiqi Zhang, Bing Zhou, Bei Wang

Traffic forecasting is essential for the traffic construction of smart cities in the new era. However, traffic data's complex spatial and temporal dependencies make traffic forecasting extremely challenging. Most existin…

ManagementTime Series Analysis

SOAP: Enhancing Spatio-Temporal Relation and Motion Information Capturing for Few-Shot Action Recognition

2024-07-23 · Wenbo Huang, Jinghui Zhang, Xuwei Qian, Zhen Wu 외

High frame-rate (HFR) videos of action recognition improve fine-grained expression while reducing the spatio-temporal relation and motion information density. Thus, large amounts of video samples are continuously require…

Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionRelation