paper-with-me

홈 › Papers

MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation

2025-08-18 · Wei Wei, Shaojie Zhang, Yonghao Dang, Jianqin Yin arxiv

Human action recognition is a crucial task for intelligent robotics, particularly within the context of human-robot collaboration research. In self-supervised skeleton-based action recognition, the mask-based reconstruction paradigm learns the spatial structure and motion patterns of the skeleton by masking joints and reconstructing the target from unlabeled data. However, existing methods focus on a limited set of joints and low-order motion patterns, limiting the model's ability to understand complex motion patterns. To address this issue, we introduce MaskSem, a novel semantic-guided masking method for learning 3D hybrid high-order motion representations. This novel framework leverages Grad-CAM based on relative motion to guide the masking of joints, which can be represented as the most semantically rich temporal orgions. The semantic-guided masking process can encourage the model to explore more discriminative features. Furthermore, we propose using hybrid high-order motion as the reconstruction target, enabling the model to learn multi-order motion patterns. Specifically, low-order motion velocity and high-order motion acceleration are used together as the reconstruction target. This approach offers a more comprehensive description of the dynamic motion process, enhancing the model's understanding of motion patterns. Experiments on the NTU60, NTU120, and PKU-MMD datasets show that MaskSem, combined with a vanilla transformer, improves skeleton-based action recognition, making it more suitable for applications in human-robot interaction.

📄 PDF Abstract BibTeX arXiv:2508.12948

Code (0)

등록된 구현이 없습니다.

Tasks

Action Recognition

Similar Papers 제목 키워드 기반

Masksembles for Uncertainty Estimation

2020-12-15 · CVPR 2021 1 · Nikita Durasov, Timur Bagautdinov, Pierre Baque, Pascal Fua

Deep neural networks have amply demonstrated their prowess but estimating the reliability of their predictions remains challenging. Deep Ensembles are widely considered as being one of the best methods for generating unc…

Classifier calibrationEnsemble LearningOut-of-Distribution DetectionRobust classification+1

SemMAE: Semantic-Guided Masking for Learning Masked Autoencoders

2022-06-21 · Gang Li, Heliang Zheng, Daqing Liu, Chaoyue Wang 외

Recently, significant progress has been made in masked image modeling to catch up to masked language modeling. However, unlike words in NLP, the lack of semantic decomposition of images still makes masked autoencoding (M…

Language ModelingLanguage ModellingMasked Language ModelingSemantic Segmentation

Semantics-Guided Multimodal Masked Autoencoder Pretraining for 3D BEV Object Detection

2026-05-24 · Prabuddhi Wariyapperuma, Rajitha de Silva, Marc Hanheide, Thomas Bohné 외 arxiv

Accurate 3D bird's-eye view (BEV) object detection is essential for autonomous driving, and depends strongly on effective multimodal representations from complementary sensors such as cameras and LiDAR. Multimodal masked…

3D Object DetectionAutonomous Driving

Self-distilled Masked Attention guided masked image modeling with noise Regularized Teacher (SMART) for medical image analysis

2023-10-02 · Jue Jiang, Aneesh Rangnekar, Chloe Min Seo Choi, Harini Veeraraghavan

Pretraining vision transformers (ViT) with attention guided masked image modeling (MIM) has shown to increase downstream accuracy for natural image analysis. Hierarchical shifted window (Swin) transformer, often used in …

Computed Tomography (CT)Medical Image Analysis

Rhamba: Region-Aware Hybrid Attention-Mamba Framework for Self-Supervised Learning in Resting-State fMRI

2026-05-02 · Ruthwik Reddy Doodipala, Pankaj Pandey, Pratheek Eranki, Carolina Torres-Rojas 외 arxiv

Self-supervised pretraining is promising for large-scale neuroimaging, yet the impact of region-aware masking and hybrid sequence modeling remains underexplored. In this work, we introduce Rhamba, a region-aware pretrain…

Self-Supervised LearningRepresentation Learning