paper-with-me

홈 › Papers

Context-Aware Network Based on Multi-scale Spatio-temporal Attention for Action Recognition in Videos

2025-12-21 · Xiaoyang Li, Wenzhu Yang, Kanglin Wang, Tiebiao Wang, Qingsong Fei arxiv

Action recognition is a critical task in video understanding, requiring the comprehensive capture of spatio-temporal cues across various scales. However, existing methods often overlook the multi-granularity nature of actions. To address this limitation, we introduce the Context-Aware Network (CAN). CAN consists of two core modules: the Multi-scale Temporal Cue Module (MTCM) and the Group Spatial Cue Module (GSCM). MTCM effectively extracts temporal cues at multiple scales, capturing both fast-changing motion details and overall action flow. GSCM, on the other hand, extracts spatial cues at different scales by grouping feature maps and applying specialized extraction methods to each group. Experiments conducted on five benchmark datasets (Something-Something V1 and V2, Diving48, Kinetics-400, and UCF101) demonstrate the effectiveness of CAN. Our approach achieves competitive performance, outperforming most mainstream methods, with accuracies of 50.4% on Something-Something V1, 63.9% on Something-Something V2, 88.4% on Diving48, 74.9% on Kinetics-400, and 86.9% on UCF101. These results highlight the importance of capturing multi-scale spatio-temporal cues for robust action recognition.

📄 PDF Abstract BibTeX arXiv:2512.18750

Code (0)

등록된 구현이 없습니다.

Tasks

Action Recognition In Videos

Similar Papers 제목 키워드 기반

Automatic ultrasound vessel segmentation with deep spatiotemporal context learning

2021-11-03 · Baichuan Jiang, Alvin Chen, Shyam Bharat, Mingxin Zheng

Accurate, real-time segmentation of vessel structures in ultrasound image sequences can aid in the measurement of lumen diameters and assessment of vascular diseases. This, however, remains a challenging task, particular…

Segmentation

Pruning for Generalization: A Transfer-Oriented Spatiotemporal Graph Framework

2026-02-04 · Zihao Jing, Yuxi Long, Ganlin Feng arxiv

Multivariate time series forecasting in graph-structured domains is critical for real-world applications, yet existing spatiotemporal models often suffer from performance degradation under data scarcity and cross-domain …

Multivariate Time Series Forecasting

ADS-POI: Agentic Spatiotemporal State Decomposition for Next Point-of-Interest Recommendation

2026-02-10 · Zhenyu Yu, Chunlei Meng, Yangchen Zeng, Mohd Yamani Idna Idris 외 arxiv

Next point-of-interest (POI) recommendation requires modeling user mobility as a spatiotemporal sequence, where different behavioral factors may evolve at different temporal and spatial scales. Most existing methods comp…

MSF-Mamba: Motion-aware State Fusion Mamba for Efficient Micro-Gesture Recognition

2025-10-12 · Deng Li, Jun Shao, Bohao Xing, Rong Gao 외 arxiv

Micro-gesture recognition (MGR) targets the identification of subtle and fine-grained human motions and requires accurate modeling of both long-range and local spatiotemporal dependencies. While CNNs are effective at cap…

Micro-gesture Recognition

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG

2026-04-07 · Honghao Fu, Miao Xu, Yiwei Wang, Dailing Zhang 외 arxiv

Scaling multimodal large language models (MLLMs) to long videos is constrained by limited context windows. While retrieval-augmented generation (RAG) is a promising remedy by organizing query-relevant visual evidence int…