paper-with-me

홈 › Papers

Temporal-Aware Self-Supervised Learning for 3D Hand Pose and Mesh Estimation in Videos

2020-12-06 · Liangjian Chen, Shih-Yao Lin, Yusheng Xie, Yen-Yu Lin, Xiaohui Xie

Estimating 3D hand pose directly from RGB imagesis challenging but has gained steady progress recently bytraining deep models with annotated 3D poses. Howeverannotating 3D poses is difficult and as such only a few 3Dhand pose datasets are available, all with limited samplesizes. In this study, we propose a new framework of training3D pose estimation models from RGB images without usingexplicit 3D annotations, i.e., trained with only 2D informa-tion. Our framework is motivated by two observations: 1)Videos provide richer information for estimating 3D posesas opposed to static images; 2) Estimated 3D poses oughtto be consistent whether the videos are viewed in the for-ward order or reverse order. We leverage these two obser-vations to develop a self-supervised learning model calledtemporal-aware self-supervised network (TASSN). By en-forcing temporal consistency constraints, TASSN learns 3Dhand poses and meshes from videos with only 2D keypointposition annotations. Experiments show that our modelachieves surprisingly good results, with 3D estimation ac-curacy on par with the state-of-the-art models trained with3D annotations, highlighting the benefit of the temporalconsistency in constraining 3D prediction models.

📄 PDF Abstract BibTeX arXiv:2012.03205

Code (0)

등록된 구현이 없습니다.

Tasks

Pose EstimationSelf-Supervised Learning

Similar Papers 제목 키워드 기반

UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimation

2026-05-18 · Tianhao Han, Haoyang Zhang, Liang Xie, Haochen Chang 외 arxiv

Manually annotating accurate 3D hand poses is extremely time-consuming and labor-intensive. Existing self-supervised hand pose estimation methods leverage the discrepancy between input images and rendered outputs, or mul…

Self-Supervised LearningHand Pose Estimation

Self-Supervised Video Object Segmentation by Motion-Aware Mask Propagation

2021-07-27 · Bo Miao, Mohammed Bennamoun, Yongsheng Gao, Ajmal Mian

We propose a self-supervised spatio-temporal matching method, coined Motion-Aware Mask Propagation (MAMP), for video object segmentation. MAMP leverages the frame reconstruction task for training without the need for ann…

SegmentationSemantic SegmentationSemi-Supervised Video Object SegmentationVideo Object Segmentation+1

SignBERT: Pre-Training of Hand-Model-Aware Representation for Sign Language Recognition

2021-10-11 · ICCV 2021 10 · Hezhen Hu, Weichao Zhao, Wengang Zhou, Yuechen Wang 외

Hand gesture serves as a critical role in sign language. Current deep-learning-based sign language recognition (SLR) methods may suffer insufficient interpretability and overfitting due to limited sign data sources. In t…

Self-Supervised LearningSign Language Recognition

SignBERT+: Hand-model-aware Self-supervised Pre-training for Sign Language Understanding

2023-05-08 · Hezhen Hu, Weichao Zhao, Wengang Zhou, Houqiang Li

Hand gesture serves as a crucial role during the expression of sign language. Current deep learning based methods for sign language understanding (SLU) are prone to over-fitting due to insufficient sign data resource and…

Self-Supervised LearningSign Language RecognitionSign Language Translation

Cross-Temporal Attention Fusion (CTAF) for Multimodal Physiological Signals in Self-Supervised Learning

2026-02-02 · Arian Khorasani, Théophile Demazure arxiv

We study multimodal affect modeling when EEG and peripheral physiology are asynchronous, which most fusion methods ignore or handle with costly warping. We propose Cross-Temporal Attention Fusion (CTAF), a self-supervise…

Self-Supervised Learning