paper-with-me

Papers

Effective Action Recognition with Embedded Key Point Shifts

2020-08-26 · Haozhi Cao, Yuecong Xu, Jianfei Yang, Kezhi Mao, Jianxiong Yin, Simon See

Temporal feature extraction is an essential technique in video-based action recognition. Key points have been utilized in skeleton-based action recognition methods but they require costly key point annotation. In this paper, we propose a novel temporal feature extraction module, named Key Point Shifts Embedding Module ($KPSEM$), to adaptively extract channel-wise key point shifts across video frames without key point annotation for temporal feature extraction. Key points are adaptively extracted as feature points with maximum feature values at split regions, while key point shifts are the spatial displacements of corresponding key points. The key point shifts are encoded as the overall temporal features via linear embedding layers in a multi-set manner. Our method achieves competitive performance through embedding key point shifts with trivial computational cost, achieving the state-of-the-art performance of 82.05% on Mini-Kinetics and competitive performance on UCF101, Something-Something-v1, and HMDB51 datasets.

📄 PDF Abstract BibTeX arXiv:2008.11378

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionSkeleton Based Action Recognition

Similar Papers 제목 키워드 기반

Video Test-Time Adaptation for Action Recognition

2022-11-24 · CVPR 2023 1 · Wei Lin, Muhammad Jehanzeb Mirza, Mateusz Kozinski, Horst Possegger 외

Although action recognition systems can achieve top performance when evaluated on in-distribution test points, they are vulnerable to unanticipated distribution shifts in test data. However, test-time adaptation of video…

Action RecognitionTemporal Action LocalizationTest-time Adaptation

Words are Malleable: Computing Semantic Shifts in Political and Media Discourse

2017-11-15 · Hosein Azarbonyad, Mostafa Dehghani, Kaspar Beelen, Alexandra Arkut 외

Recently, researchers started to pay attention to the detection of temporal shifts in the meaning of words. However, most (if not all) of these approaches restricted their efforts to uncovering change over time, thus neg…

Text-Embedded Bilinear Model for Fine-Grained Visual Recognition

2020-10-12 · Liang Sun, Xiang Guan, Yang Yang, Lei Zhang

Fine-grained visual recognition, which aims to identify subcategories of the same base-level category, is a challenging task because of its large intra-class variances and small inter-class variances. Human beings can p…

Fine-Grained Image RecognitionFine-Grained Visual RecognitionObject Recognition

EV-CLIP: Efficient Visual Prompt Adaptation for CLIP in Few-shot Action Recognition under Visual Challenges

2026-04-24 · Hyo Jin Jon, Longbin Jin, Eun Yi Kim arxiv

CLIP has demonstrated strong generalization in visual domains through natural language supervision, even for video action recognition. However, most existing approaches that adapt CLIP for action recognition have primari…

Action Recognition

Cross-Domain Human Action Recognition from Multiview Motion and Textual Descriptions

2026-05-21 · Yannick Porto, Renato Martins, Thomas Chalumeau, Cedric Demonceaux arxiv

Robustness to domain changes is a key capability for effective deployment of human action recognition systems in real-world scenarios, where action categories at inference can present important domain shifts or even unse…

Zero-Shot Action RecognitionTransfer Learning