paper-with-me

Papers

Spatio-temporal Self-Supervised Representation Learning for 3D Point Clouds

2021-09-01 · ICCV 2021 10 · Siyuan Huang, Yichen Xie, Song-Chun Zhu, Yixin Zhu

To date, various 3D scene understanding tasks still lack practical and generalizable pre-trained models, primarily due to the intricate nature of 3D scene understanding tasks and their immense variations introduced by camera views, lighting, occlusions, etc. In this paper, we tackle this challenge by introducing a spatio-temporal representation learning (STRL) framework, capable of learning from unlabeled 3D point clouds in a self-supervised fashion. Inspired by how infants learn from visual data in the wild, we explore the rich spatio-temporal cues derived from the 3D data. Specifically, STRL takes two temporally-correlated frames from a 3D point cloud sequence as the input, transforms it with the spatial data augmentation, and learns the invariant representation self-supervisedly. To corroborate the efficacy of STRL, we conduct extensive experiments on three types (synthetic, indoor, and outdoor) of datasets. Experimental results demonstrate that, compared with supervised learning methods, the learned self-supervised representation facilitates various models to attain comparable or even better performances while capable of generalizing pre-trained models to downstream tasks, including 3D shape classification, 3D object detection, and 3D semantic segmentation. Moreover, the spatio-temporal contextual cues embedded in 3D point clouds significantly improve the learned representations.

📄 PDF Abstract BibTeX arXiv:2109.00179

Code (1)

yichen928/STRL 공식 구현 pytorch

Tasks

3D Object Detection3D Point Cloud Classification3D Point Cloud Linear Classification3D Semantic Segmentation3D Shape ClassificationData Augmentationobject-detectionObject DetectionRepresentation LearningScene UnderstandingSemantic SegmentationUnsupervised 3D Point Cloud Linear Evaluation

Similar Papers 제목 키워드 기반

Video Playback Rate Perception for Self-supervisedSpatio-Temporal Representation Learning

2020-06-20 · Yuan Yao, Chang Liu, Dezhao Luo, Yu Zhou 외

In self-supervised spatio-temporal representation learning, the temporal resolution and long-short term characteristics are not yet fully explored, which limits representation capabilities of learned models. In this pape…

Action RecognitionDecoderRepresentation LearningRetrieval+1

Video Playback Rate Perception for Self-Supervised Spatio-Temporal Representation Learning

2020-06-01 · CVPR 2020 6 · Yuan Yao, Chang Liu, Dezhao Luo, Yu Zhou 외

In self-supervised spatio-temporal representation learning, the temporal resolution and long-short term characteristics are not yet fully explored, which limits representation capabilities of learned models. In this pape…

Action RecognitionDecoderRepresentation LearningRetrieval+3

Masked Spatio-Temporal Structure Prediction for Self-supervised Learning on Point Cloud Videos

2023-08-18 · ICCV 2023 1 · Zhiqiang Shen, Xiaoxiao Sheng, Hehe Fan, Longguang Wang 외

Recently, the community has made tremendous progress in developing effective methods for point cloud video understanding that learn from massive amounts of labeled data. However, annotating point cloud videos is usually …

point cloud video understandingSelf-Supervised LearningVideo Understanding

Self-Supervised Video Representation Learning with Constrained Spatiotemporal Jigsaw

2021-01-01 · Yuqi Huo, Mingyu Ding, Haoyu Lu, Zhiwu Lu 외

This paper proposes a novel pretext task for self-supervised video representation learning by exploiting spatiotemporal continuity in videos. It is motivated by the fact that videos are spatiotemporal by nature and a rep…

Representation Learning

Two Stream Self-Supervised Learning for Action Recognition

2018-06-16 · Ahmed Taha, Moustafa Meshry, Xitong Yang, Yi-Ting Chen 외

We present a self-supervised approach using spatio-temporal signals between video frames for action recognition. A two-stream architecture is leveraged to tangle spatial and temporal representation learning. Our task is …

Action RecognitionRepresentation LearningSelf-Supervised LearningTemporal Action Localization+1