SignBERT: Pre-Training of Hand-Model-Aware Representation for Sign Language Recognition
Hand gesture serves as a critical role in sign language. Current deep-learning-based sign language recognition (SLR) methods may suffer insufficient interpretability and overfitting due to limited sign data sources. In this paper, we introduce the first self-supervised pre-trainable SignBERT with incorporated hand prior for SLR. SignBERT views the hand pose as a visual token, which is derived from an off-the-shelf pose extractor. The visual tokens are then embedded with gesture state, temporal and hand chirality information. To take full advantage of available sign data sources, SignBERT first performs self-supervised pre-training by masking and reconstructing visual tokens. Jointly with several mask modeling strategies, we attempt to incorporate hand prior in a model-aware method to better model hierarchical context over the hand sequence. Then with the prediction head added, SignBERT is fine-tuned to perform the downstream SLR task. To validate the effectiveness of our method on SLR, we perform extensive experiments on four public benchmark datasets, i.e., NMFs-CSL, SLR500, MSASL and WLASL. Experiment results demonstrate the effectiveness of both self-supervised learning and imported hand prior. Furthermore, we achieve state-of-the-art performance on all benchmarks with a notable gain.
Code (0)
등록된 구현이 없습니다.
Tasks
Self-Supervised LearningSign Language RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
SignBERT+: Hand-model-aware Self-supervised Pre-training for Sign Language Understanding
Hand gesture serves as a crucial role during the expression of sign language. Current deep learning based methods for sign language understanding (SLU) are prone to over-fitting due to insufficient sign data resource and…
Self-Supervised LearningSign Language RecognitionSign Language TranslationGesture-Aware Pretraining and Token Fusion for 3D Hand Pose Estimation
Estimating 3D hand pose from monocular RGB images is fundamental for applications in AR/VR, human-computer interaction, and sign language understanding. In this work we focus on a scenario where a discrete set of gesture…
3D Hand Pose Estimation3D Pose EstimationGeometry-Aware Self-Training for Unsupervised Domain Adaptationon Object Point Clouds
The point cloud representation of an object can have a large geometric variation in view of inconsistent data acquisition procedure, which thus leads to domain discrepancy due to diverse and uncontrollable shape represen…
Domain AdaptationPoint Cloud ClassificationRepresentation LearningUnsupervised Domain AdaptationGeometry-Aware Self-Training for Unsupervised Domain Adaptation on Object Point Clouds
The point cloud representation of an object can have a large geometric variation in view of inconsistent data acquisition procedure, which thus leads to domain discrepancy due to diverse and uncontrollable shape repr…
Domain AdaptationPoint Cloud ClassificationRepresentation LearningUnsupervised Domain AdaptationCLIP-Hand3D: Exploiting 3D Hand Pose Estimation via Context-Aware Prompting
Contrastive Language-Image Pre-training (CLIP) starts to emerge in many computer vision tasks and has achieved promising performance. However, it remains underexplored whether CLIP can be generalized to 3D hand pose esti…
3D Hand Pose EstimationContrastive LearningHand Pose EstimationPose Estimation