paper-with-me

홈 › Papers

Attention Transfer from Web Images for Video Recognition

2017-08-03 · Junnan Li, Yongkang Wong, Qi Zhao, Mohan Kankanhalli

Training deep learning based video classifiers for action recognition requires a large amount of labeled videos. The labeling process is labor-intensive and time-consuming. On the other hand, large amount of weakly-labeled images are uploaded to the Internet by users everyday. To harness the rich and highly diverse set of Web images, a scalable approach is to crawl these images to train deep learning based classifier, such as Convolutional Neural Networks (CNN). However, due to the domain shift problem, the performance of Web images trained deep classifiers tend to degrade when directly deployed to videos. One way to address this problem is to fine-tune the trained models on videos, but sufficient amount of annotated videos are still required. In this work, we propose a novel approach to transfer knowledge from image domain to video domain. The proposed method can adapt to the target domain (i.e. video data) with limited amount of training data. Our method maps the video frames into a low-dimensional feature space using the class-discriminative spatial attention map for CNNs. We design a novel Siamese EnergyNet structure to learn energy functions on the attention maps by jointly optimizing two loss functions, such that the attention map corresponding to a ground truth concept would have higher energy. We conduct extensive experiments on two challenging video recognition datasets (i.e. TVHI and UCF101), and demonstrate the efficacy of our proposed method.

📄 PDF Abstract BibTeX arXiv:1708.00973

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionTemporal Action LocalizationVideo Recognition

Similar Papers 제목 키워드 기반

Red Carpet to Fight Club: Partially-supervised Domain Transfer for Face Recognition in Violent Videos

2020-09-16 · Yunus Can Bilge, Mehmet Kerim Yucel, Ramazan Gokberk Cinbis, Nazli Ikizler-Cinbis 외

In many real-world problems, there is typically a large discrepancy between the characteristics of data used in training versus deployment. A prime example is the analysis of aggression videos: in a criminal incidence, t…

Face Recognition

ZeroI2V: Zero-Cost Adaptation of Pre-trained Transformers from Image to Video

2023-10-02 · Xinhao Li, Yuhan Zhu, LiMin Wang

Adapting image models to the video domain has emerged as an efficient paradigm for solving video recognition tasks. Due to the huge number of parameters and effective transferability of image models, performing full fine…

Action ClassificationAction RecognitionVideo Recognition

Extreme Low Resolution Activity Recognition with Confident Spatial-Temporal Attention Transfer

2019-09-09 · Yucai Bai, Qin Zou, Xieyuanli Chen, Lingxi Li 외

Activity recognition on extreme low-resolution videos, e.g., a resolution of 12*16 pixels, plays a vital role in far-view surveillance and privacy-preserving multimedia analysis. Low-resolution videos only contain limite…

Activity RecognitionPrivacy PreservingTransfer Learning

Cross-view Action Recognition Understanding From Exocentric to Egocentric Perspective

2023-05-25 · Thanh-Dat Truong, Khoa Luu

Understanding action recognition in egocentric videos has emerged as a vital research topic with numerous practical applications. With the limitation in the scale of egocentric data collection, learning robust deep learn…

Action Recognition

CTrGAN: Cycle Transformers GAN for Gait Transfer

2022-06-30 · Shahar Mahpod, Noam Gaash, Hay Hoffman, Gil Ben-Artzi

We introduce a novel approach for gait transfer from unconstrained videos in-the-wild. In contrast to motion transfer, the objective here is not to imitate the source's motions by the target, but rather to replace the wa…

DecoderGait Recognition