paper-with-me

Papers

Motion Capture from Internet Videos

2020-08-18 · ECCV 2020 8 · Junting Dong, Qing Shuai, Yuanqing Zhang, Xian Liu, Xiaowei Zhou, Hujun Bao

Recent advances in image-based human pose estimation make it possible to capture 3D human motion from a single RGB video. However, the inherent depth ambiguity and self-occlusion in a single view prohibit the recovery of as high-quality motion as multi-view reconstruction. While multi-view videos are not common, the videos of a celebrity performing a specific action are usually abundant on the Internet. Even if these videos were recorded at different time instances, they would encode the same motion characteristics of the person. Therefore, we propose to capture human motion by jointly analyzing these Internet videos instead of using single videos separately. However, this new task poses many new challenges that cannot be addressed by existing methods, as the videos are unsynchronized, the camera viewpoints are unknown, the background scenes are different, and the human motions are not exactly the same among videos. To address these challenges, we propose a novel optimization-based framework and experimentally demonstrate its ability to recover much more precise and detailed motion from multiple videos, compared against monocular motion capture methods.

📄 PDF Abstract BibTeX arXiv:2008.07931

Code (2)

zju3dv/iMoCap 공식 구현 pytorch
zju3dv/EasyMocap pytorch

Tasks

Pose Estimation

Similar Papers 제목 키워드 기반

AnyLift: Scaling Motion Reconstruction from Internet Videos via 2D Diffusion

2026-04-20 · Hongjie Li, Heng Yu, Jiaman Li, Hong-Xing Yu 외 arxiv

Reconstructing 3D human motion and human-object interactions (HOI) from Internet videos is a fundamental step toward building large-scale datasets of human behavior. Existing methods struggle to recover globally consiste…

PHUMA: Physically Reliable Humanoid Locomotion Dataset

2025-10-30 · Kyungmin Lee, Sibeen Kim, Youngdo Lee, Minho Park 외 arxiv

Motion imitation is a promising approach for humanoid locomotion, enabling agents to acquire humanlike behaviors. Existing methods typically rely on high-quality motion capture datasets such as AMASS, but these are scarc…

Modelling Temporal Information Using Discrete Fourier Transform for Recognizing Emotions in User-generated Videos

2016-03-20 · Haimin Zhang, Min Xu

With the widespread of user-generated Internet videos, emotion recognition in those videos attracts increasing research efforts. However, most existing works are based on framelevel visual features and/or audio features,…

Emotion ClassificationEmotion RecognitionVideo Emotion Recognition

What Matters When Cotraining Robot Manipulation Policies on Everyday Human Videos?

2026-06-04 · Richard Li, Aditya Prakash, Andrew Wen, Saurabh Gupta 외 arxiv

Human video datasets used for cotraining robot manipulation policies largely consist of curated demonstrations where motions are orchestrated to resemble robot behavior and 3D hand poses are captured with specialized har…

Robot Manipulation

CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning

2025-05-22 · Jiange Yang, Yansong Shi, Haoyi Zhu, MingYu Liu 외

Learning latent motion from Internet videos is crucial for building generalist robots. However, existing discrete latent action methods suffer from information loss and struggle with complex and fine-grained dynamics. We…

Zero-shot Generalization