paper-with-me

Papers

TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding

2023-10-29 · Shuhuai Ren, Sishuo Chen, Shicheng Li, Xu sun, Lu Hou

Large-scale video-language pre-training has made remarkable strides in advancing video-language understanding tasks. However, the heavy computational burden of video encoding remains a formidable efficiency bottleneck, particularly for long-form videos. These videos contain massive visual tokens due to their inherent 3D properties and spatiotemporal redundancy, making it challenging to capture complex temporal and spatial relationships. To tackle this issue, we propose an efficient method called TEmporal-Spatial Token Aggregation (TESTA). TESTA condenses video semantics by adaptively aggregating similar frames, as well as similar patches within each frame. TESTA can reduce the number of visual tokens by 75% and thus accelerate video encoding. Building upon TESTA, we introduce a pre-trained video-language model equipped with a divided space-time token aggregation module in each video encoder block. We evaluate our model on five datasets for paragraph-to-video retrieval and long-form VideoQA tasks. Experimental results show that TESTA improves computing efficiency by 1.7 times, and achieves significant performance gains from its scalability in processing longer input frames, e.g., +13.7 R@1 on QuerYD and +6.5 R@1 on Condensed Movie.

📄 PDF Abstract BibTeX arXiv:2310.19060

Code (1)

renshuhuai-andy/testa 공식 구현 pytorch

Tasks

FormLanguage ModellingRetrievalVideo Question AnsweringVideo RetrievalVideo-Text Retrieval

Similar Papers 제목 키워드 기반

Less is More: Strategic Expert Selection Outperforms Ensemble Complexity in Traffic Forecasting

2025-10-08 · Walid Guettala, Yufan Zhao, László Gulyás arxiv

Traffic forecasting is fundamental to intelligent transportation systems, enabling congestion mitigation and emission reduction in increasingly complex urban environments. While recent graph neural network approaches hav…

Computational EfficiencyGraph Neural Network

TESTAM: A Time-Enhanced Spatio-Temporal Attention Model with Mixture of Experts

2024-03-05 · Hyunwook Lee, Sungahn Ko

Accurate traffic forecasting is challenging due to the complex dependency on road networks, various types of roads, and the abrupt speed change due to the events. Recent works mainly focus on dynamic spatial modeling wit…

Graph AttentionGraph EmbeddingMixture-of-Experts

STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation

2026-07-03 · Syed Ariff Syed Hesham, Yun Liu, Guolei Sun, Jing Yang 외 arxiv

Video reasoning segmentation demands pixel-accurate object tracking across hundreds of frames under complex natural language queries, producing dense spatiotemporal tokens whose quadratic self-attention cost makes long-v…

Natural Language QueriesObject Tracking

Generative Video Compression with One-Dimensional Latent Representation

2026-03-16 · Zihan Zheng, Zhaoyang Jia, Naifu Xue, Jiahao Li 외 arxiv

Recent advancements in generative video codec (GVC) typically encode video into a 2D latent grid and employ high-capacity generative decoders for reconstruction. However, this paradigm still leaves two key challenges in …

Trajectory-aware Shifted State Space Models for Online Video Super-Resolution

2025-08-14 · Qiang Zhu, Xiandong Meng, Yuxian Jiang, Fan Zhang 외 arxiv

Online video super-resolution (VSR) is an important technique for many real-world video processing applications, which aims to restore the current high-resolution video frame based on temporally previous frames. Most of …

Computational EfficiencyVideo Super-ResolutionTrajectory Modeling