paper-with-me

Papers

Unsupervised Spatial-Temporal Feature Enrichment and Fidelity Preservation Network for Skeleton based Action Recognition

2024-01-25 · Chuankun Li, Shuai Li, Yanbo Gao, Ping Chen, Jian Li, Wanqing Li

Unsupervised skeleton based action recognition has achieved remarkable progress recently. Existing unsupervised learning methods suffer from severe overfitting problem, and thus small networks are used, significantly reducing the representation capability. To address this problem, the overfitting mechanism behind the unsupervised learning for skeleton based action recognition is first investigated. It is observed that the skeleton is already a relatively high-level and low-dimension feature, but not in the same manifold as the features for action recognition. Simply applying the existing unsupervised learning method may tend to produce features that discriminate the different samples instead of action classes, resulting in the overfitting problem. To solve this problem, this paper presents an Unsupervised spatial-temporal Feature Enrichment and Fidelity Preservation framework (U-FEFP) to generate rich distributed features that contain all the information of the skeleton sequence. A spatial-temporal feature transformation subnetwork is developed using spatial-temporal graph convolutional network and graph convolutional gate recurrent unit network as the basic feature extraction network. The unsupervised Bootstrap Your Own Latent based learning is used to generate rich distributed features and the unsupervised pretext task based learning is used to preserve the information of the skeleton sequence. The two unsupervised learning ways are collaborated as U-FEFP to produce robust and discriminative representations. Experimental results on three widely used benchmarks, namely NTU-RGB+D-60, NTU-RGB+D-120 and PKU-MMD dataset, demonstrate that the proposed U-FEFP achieves the best performance compared with the state-of-the-art unsupervised learning methods. t-SNE illustrations further validate that U-FEFP can learn more discriminative features for unsupervised skeleton based action recognition.

📄 PDF Abstract BibTeX arXiv:2401.14034

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionSkeleton Based Action RecognitionUnsupervised Skeleton Based Action Recognition

Similar Papers 제목 키워드 기반

Spatio-temporal Relation Modeling for Few-shot Action Recognition

2021-12-09 · CVPR 2022 1 · Anirudh Thatipelli, Sanath Narayan, Salman Khan, Rao Muhammad Anwer 외

We propose a novel few-shot action recognition framework, STRM, which enhances class-specific feature discriminability while simultaneously learning higher-order temporal representations. The focus of our approach is a n…

Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionRelation

ContextFlow: Training-Free Video Object Editing via Adaptive Context Enrichment

2025-09-22 · Yiyang Chen, Xuanhua He, Xiujun Ma, Yue Ma arxiv

Training-free video object editing aims to achieve precise object-level manipulation, including object insertion, swapping, and deletion. However, it faces significant challenges in maintaining fidelity and temporal cons…

Unsupervised Spiking Instance Segmentation on Event Data using STDP

2021-11-09 · Paul Kirkland, Davide L. Manna, Alex Vicente-Sola, Gaetano Di Caterina

Spiking Neural Networks (SNN) and the field of Neuromorphic Engineering has brought about a paradigm shift in how to approach Machine Learning (ML) and Computer Vision (CV) problem. This paradigm shift comes from the ada…

Event-based visionFace DetectionFace RecognitionInstance Segmentation+3

Enriched Long-term Recurrent Convolutional Network for Facial Micro-Expression Recognition

2018-05-22 · Huai-Qian Khor, John See, Raphael C. -W. Phan, Weiyao Lin

Facial micro-expression (ME) recognition has posed a huge challenge to researchers for its subtlety in motion and limited databases. Recently, handcrafted techniques have achieved superior performance in micro-expression…

Data AugmentationMicro Expression RecognitionMicro-Expression RecognitionSpecificity

CAE-AV: Improving Audio-Visual Learning via Cross-modal Interactive Enrichment

2026-02-09 · Yunzuo Hu, Wen Li, Jing Zhang arxiv

Audio-visual learning suffers from modality misalignment caused by off-screen sources and background clutter, and current methods usually amplify irrelevant regions or moments, leading to unstable training and degraded r…