paper-with-me

Papers

Hierarchically Decoupled Spatial-Temporal Contrast for Self-supervised Video Representation Learning

2020-11-23 · Zehua Zhang, David Crandall

We present a novel technique for self-supervised video representation learning by: (a) decoupling the learning objective into two contrastive subtasks respectively emphasizing spatial and temporal features, and (b) performing it hierarchically to encourage multi-scale understanding. Motivated by their effectiveness in supervised learning, we first introduce spatial-temporal feature learning decoupling and hierarchical learning to the context of unsupervised video learning. We show by experiments that augmentations can be manipulated as regularization to guide the network to learn desired semantics in contrastive learning, and we propose a way for the model to separately capture spatial and temporal features at multiple scales. We also introduce an approach to overcome the problem of divergent levels of instance invariance at different hierarchies by modeling the invariance as loss weights for objective re-weighting. Experiments on downstream action recognition benchmarks on UCF101 and HMDB51 show that our proposed Hierarchically Decoupled Spatial-Temporal Contrast (HDC) makes substantial improvements over directly learning spatial-temporal features as a whole and achieves competitive performance when compared with other state-of-the-art unsupervised methods. Code will be made available.

📄 PDF Abstract BibTeX arXiv:2011.11261

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionContrastive LearningRepresentation Learning

Similar Papers 제목 키워드 기반

Hierarchically Self-Supervised Transformer for Human Skeleton Representation Learning

2022-07-20 · Yuxiao Chen, Long Zhao, Jianbo Yuan, Yu Tian 외

Despite the success of fully-supervised human skeleton sequence modeling, utilizing self-supervised pre-training for skeleton sequence representation learning has been an active field because acquiring task-specific skel…

Action DetectionAction RecognitionContrastive Learningmotion prediction+1

Decoupled Spatio-Temporal Consistency Learning for Self-Supervised Tracking

2025-07-29 · Yaozong Zheng, Bineng Zhong, Qihua Liang, Ning Li 외 arxiv

The success of visual tracking has been largely driven by datasets with manual box annotations. However, these box annotations require tremendous human effort, limiting the scale and diversity of existing tracking datase…

Visual Tracking

SDMTL: Semi-Decoupled Multi-grained Trajectory Learning for 3D human motion prediction

2020-10-11 · Xiaoli Liu, Jianqin Yin

Predicting future human motion is critical for intelligent robots to interact with humans in the real world, and human motion has the nature of multi-granularity. However, most of the existing work either implicitly mode…

Human motion predictionmotion prediction

Spatial-Temporal-Decoupled Masked Pre-training for Spatiotemporal Forecasting

2023-12-01 · Haotian Gao, Renhe Jiang, Zheng Dong, Jinliang Deng 외

Spatiotemporal forecasting techniques are significant for various domains such as transportation, energy, and weather. Accurate prediction of spatiotemporal series remains challenging due to the complex spatiotemporal he…

Time SeriesTraffic Prediction

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation

2024-11-21 · Lin Sun, Jiale Cao, Jin Xie, Xiaoheng Jiang 외

Contrastive Language-Image Pre-training (CLIP) exhibits strong zero-shot classification ability on various image-level tasks, leading to the research to adapt CLIP for pixel-level open-vocabulary semantic segmentation wi…

Open Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationSegmentationSemantic Segmentation+2