paper-with-me

홈 › Papers

Exploiting Images for Video Recognition with Hierarchical Generative Adversarial Networks

2018-05-11 · Feiwu Yu, Xinxiao wu, Yuchao Sun, Lixin Duan

Existing deep learning methods of video recognition usually require a large number of labeled videos for training. But for a new task, videos are often unlabeled and it is also time-consuming and labor-intensive to annotate them. Instead of human annotation, we try to make use of existing fully labeled images to help recognize those videos. However, due to the problem of domain shifts and heterogeneous feature representations, the performance of classifiers trained on images may be dramatically degraded for video recognition tasks. In this paper, we propose a novel method, called Hierarchical Generative Adversarial Networks (HiGAN), to enhance recognition in videos (i.e., target domain) by transferring knowledge from images (i.e., source domain). The HiGAN model consists of a \emph{low-level} conditional GAN and a \emph{high-level} conditional GAN. By taking advantage of these two-level adversarial learning, our method is capable of learning a domain-invariant feature representation of source images and target videos. Comprehensive experiments on two challenging video recognition datasets (i.e. UCF101 and HMDB51) demonstrate the effectiveness of the proposed method when compared with the existing state-of-the-art domain adaptation methods.

📄 PDF Abstract BibTeX arXiv:1805.04384

Code (0)

등록된 구현이 없습니다.

Tasks

Domain AdaptationVideo Recognition

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Dogecoin Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

A Probabilistic Framework for Hierarchical Goal Recognition

2026-04-24 · Chenyuan Zhang, Katherine Ip, Hamid Rezatofighi, Buser Say 외 arxiv

Goal recognition aims to infer an agent's goal from observations of its behaviour. In realistic settings, recognition can benefit from exploiting hierarchical task structure and reasoning under uncertainty. Planning-base…

Exploiting temporal consistency for real-time video depth estimation

2019-08-10 · ICCV 2019 10 · Haokui Zhang, Chunhua Shen, Ying Li, Yuanzhouhan Cao 외

Accuracy of depth estimation from static images has been significantly improved recently, by exploiting hierarchical features from deep convolutional neural networks (CNNs). Compared with static images, vast information …

Depth EstimationMonocular Depth Estimation

Exploiting Motion Information from Unlabeled Videos for Static Image Action Recognition

2019-12-01 · Yiyi Zhang, Li Niu, Ziqi Pan, Meichao Luo 외

Static image action recognition, which aims to recognize action based on a single image, usually relies on expensive human labeling effort such as adequate labeled action images and large-scale labeled image dataset. In …

Action RecognitionSelf-Supervised Learning

Skepxels: Spatio-temporal Image Representation of Human Skeleton Joints for Action Recognition

2017-11-16 · Jian Liu, Naveed Akhtar, Ajmal Mian

Human skeleton joints are popular for action analysis since they can be easily extracted from videos to discard background noises. However, current skeleton representations do not fully benefit from machine learning with…

Action AnalysisAction RecognitionTemporal Action Localization

MirrorNet: A Deep Bayesian Approach to Reflective 2D Pose Estimation from Human Images

2020-04-08 · Takayuki Nakatsuka, Kazuyoshi Yoshii, Yuki Koyama, Satoru Fukayama 외

This paper proposes a statistical approach to 2D pose estimation from human images. The main problems with the standard supervised approach, which is based on a deep recognition (image-to-pose) model, are that it often y…

2D Pose EstimationPose Estimation