paper-with-me

홈 › Papers

Learning Video Representations of Human Motion From Synthetic Data

2022-01-01 · CVPR 2022 1 · Xi Guo, Wei Wu, Dongliang Wang, Jing Su, Haisheng Su, Weihao Gan, Jian Huang, Qin Yang

In this paper, we take an early step towards video representation learning of human actions with the help of largescale synthetic videos, particularly for human motion representation enhancement. Specifically, we first introduce an automatic action-related video synthesis pipeline based on a photorealistic video game. A large-scale human action dataset named GATA (GTA Animation Transformed Actions) is then built by the proposed pipeline, which includes 8.1 million action clips spanning over 28K action classes. Based on the presented dataset, we design a contrastive learning framework for human motion representation learning, which shows significant performance improvements on several typical video datasets for action recognition, e.g., Charades, HAA 500 and NTU-RGB. Besides, we further explore a domain adaptation method based on cross-domain positive pairs mining to alleviate the domain gap between synthetic and realistic data. Extensive properties analyses of learned representation are conducted to demonstrate the effectiveness of the proposed dataset for enhancing human motion representation learning.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionContrastive LearningDomain AdaptationRepresentation Learning

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics

2025-08-13 · Weiqi Li, Zehao Zhang, Liang Lin, Guangrun Wang arxiv

\textbf{Synthetic human dynamics} aims to generate photorealistic videos of human subjects performing expressive, intention-driven motions. However, current approaches face two core challenges: (1) \emph{geometric incons…

From Frames to Sequences: Temporally Consistent Human-Centric Dense Prediction

2026-02-02 · Xingyu Miao, Junting Dong, Qin Zhao, Yuhang Yang 외 arxiv

In this work, we focus on the challenge of temporally consistent human-centric dense prediction across video sequences. Existing models achieve strong per-frame accuracy but often flicker under motion, occlusion, and lig…

Exploring the Role of Synthetic Data Augmentation in Controllable Human-Centric Video Generation

2026-04-23 · Yuanchen Fei, Yude Zou, Zejian Kang, Ming Li 외 arxiv

Controllable human video generation aims to produce realistic videos of humans with explicitly guided motions and appearances,serving as a foundation for digital humans, animation, and embodied AI.However, the scarcity o…

Data AugmentationVideo Generation

Self-Supervised Learning of Structured Dynamics from Videos

2026-07-23 · Lukas Knobel, Andrew Zisserman, Yuki M. Asano arxiv

Understanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two sources of dynamics: camera motion and object motion. This decomposition has remained underexplored in …

Self-Supervised LearningRepresentation Learning

LocoMotion: Learning Motion-Focused Video-Language Representations

2024-10-15 · Hazel Doughty, Fida Mohammad Thoker, Cees G. M. Snoek

This paper strives for motion-focused video-language representations. Existing methods to learn video-language representations use spatial-focused data, where identifying the objects and scene is often enough to distingu…