Skeleton-Based Human Action Recognition with Noisy Labels
Understanding human actions from body poses is critical for assistive robots sharing space with humans in order to make informed and safe decisions about the next interaction. However, precise temporal localization and annotation of activity sequences is time-consuming and the resulting labels are often noisy. If not effectively addressed, label noise negatively affects the model's training, resulting in lower recognition quality. Despite its importance, addressing label noise for skeleton-based action recognition has been overlooked so far. In this study, we bridge this gap by implementing a framework that augments well-established skeleton-based human action recognition methods with label-denoising strategies from various research areas to serve as the initial benchmark. Observations reveal that these baselines yield only marginal performance when dealing with sparse skeleton data. Consequently, we introduce a novel methodology, NoiseEraSAR, which integrates global sample selection, co-teaching, and Cross-Modal Mixture-of-Experts (CM-MOE) strategies, aimed at mitigating the adverse impacts of label noise. Our proposed approach demonstrates better performance on the established benchmark, setting new state-of-the-art standards. The source code for this study is accessible at https://github.com/xuyizdby/NoiseEraSAR.
Code (1)
Tasks
Action RecognitionDenoisingMixture-of-ExpertsSkeleton Based Action RecognitionTemporal Action LocalizationTemporal LocalizationSimilar Papers 제목 키워드 기반
Predictively Encoded Graph Convolutional Network for Noise-Robust Skeleton-based Action Recognition
In skeleton-based action recognition, graph convolutional networks (GCNs), which model human body skeletons using graphical components such as nodes and connections, have achieved remarkable performance recently. However…
Action RecognitionSkeleton Based Action RecognitionEnhanced skeleton visualization for view invariant human action recognition
Human action recognition based on skeletons has wide applications in human–computer interaction and intelligent surveillance. However, view variations and noisy data bring challenges to this task. What’s more, it remains…
Action RecognitionSkeleton Based Action RecognitionTemporal Action LocalizationGenerative Data Augmentation for Skeleton Action Recognition
Skeleton-based human action recognition is a powerful approach for understanding human behaviour from pose data, but collecting large-scale, diverse, and well-annotated 3D skeleton datasets is both expensive and labor-in…
Action RecognitionData AugmentationSkeleton-Based Mutually Assisted Interacted Object Localization and Human Action Recognition
Skeleton data carries valuable motion information and is widely explored in human action recognition. However, not only the motion information but also the interaction with the environment provides discriminative cues to…
Action RecognitionObjectObject LocalizationTemporal Action LocalizationSkeleton Cloud Colorization for Unsupervised 3D Action Representation Learning
Skeleton-based human action recognition has attracted increasing attention in recent years. However, most of the existing works focus on supervised learning which requiring a large number of annotated action sequences th…
3D Action RecognitionAction RecognitionColorizationRepresentation Learning+1