paper-with-me

홈 › Papers

Learning Human Pose Models from Synthesized Data for Robust RGB-D Action Recognition

2017-07-04 · Jian Liu, Naveed Akhtar, Ajmal Mian

We propose Human Pose Models that represent RGB and depth images of human poses independent of clothing textures, backgrounds, lighting conditions, body shapes and camera viewpoints. Learning such universal models requires training images where all factors are varied for every human pose. Capturing such data is prohibitively expensive. Therefore, we develop a framework for synthesizing the training data. First, we learn representative human poses from a large corpus of real motion captured human skeleton data. Next, we fit synthetic 3D humans with different body shapes to each pose and render each from 180 camera viewpoints while randomly varying the clothing textures, background and lighting. Generative Adversarial Networks are employed to minimize the gap between synthetic and real image distributions. CNN models are then learned that transfer human poses to a shared high-level invariant space. The learned CNN models are then used as invariant feature extractors from real RGB and depth frames of human action videos and the temporal variations are modelled by Fourier Temporal Pyramid. Finally, linear SVM is used for classification. Experiments on three benchmark cross-view human action datasets show that our algorithm outperforms existing methods by significant margins for RGB only and RGB-D action recognition.

📄 PDF Abstract BibTeX arXiv:1707.00823

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionSkeleton Based Action RecognitionTemporal Action Localization

Methods 이 논문이 사용한 방법론

SVM A Support Vector Machine, or SVM, is a non-parametric supervised learning model. For non-linear classification and regression, they utilise the kernel trick to map inputs…

Similar Papers 제목 키워드 기반

Speech Recognition with Augmented Synthesized Speech

2019-09-25 · Andrew Rosenberg, Yu Zhang, Bhuvana Ramabhadran, Ye Jia 외

Recent success of the Tacotron speech synthesis architecture and its variants in producing natural sounding multi-speaker synthesized speech has raised the exciting possibility of replacing expensive, manually transcribe…

Data AugmentationDiversityRobust Speech Recognitionspeech-recognition+2

Action-Conditioned 3D Human Motion Synthesis with Transformer VAE

2021-04-12 · ICCV 2021 10 · Mathis Petrovich, Michael J. Black, Gül Varol

We tackle the problem of action-conditioned generation of realistic and diverse human motion sequences. In contrast to methods that complete, or extend, motion sequences, this task does not require an initial pose or seq…

Action RecognitionDenoisingHuman action generationMotion Synthesis

On the Emotion Understanding of Synthesized Speech

2026-03-17 · Yuan Ge, Haishu Zhao, Aokai Hao, Junxiang Zhang 외 arxiv

Emotion is a core paralinguistic feature in voice interaction. It is widely believed that emotion understanding models learn fundamental representations that transfer to synthesized speech, making emotion understanding r…

Speech Emotion RecognitionSpeech Synthesis

Synthesized Annotation Guidelines are Knowledge-Lite Boosters for Clinical Information Extraction

2025-04-01 · Enshuo Hsu, Martin Ugbala, Krishna Kumar Kookal, Zouaidi Kawtar 외

Generative information extraction using large language models, particularly through few-shot learning, has become a popular method. Recent studies indicate that providing a detailed, human-readable guideline-similar to t…

Few-Shot Learningnamed-entity-recognitionNamed Entity RecognitionText Generation

Thermal to Visible Synthesis of Face Images using Multiple Regions

2018-03-20 · Benjamin S. Riggan, Nathaniel J. Short, Shuowen Hu

Synthesis of visible spectrum faces from thermal facial imagery is a promising approach for heterogeneous face recognition; enabling existing face recognition software trained on visible imagery to be leveraged, and allo…

Face RecognitionFacial Landmark DetectionHeterogeneous Face RecognitionImage Registration