paper-with-me

Papers

ViewCLR: Learning Self-supervised Video Representation for Unseen Viewpoints

2021-12-07 · Srijan Das, Michael S. Ryoo

Learning self-supervised video representation predominantly focuses on discriminating instances generated from simple data augmentation schemes. However, the learned representation often fails to generalize over unseen camera viewpoints. To this end, we propose ViewCLR, that learns self-supervised video representation invariant to camera viewpoint changes. We introduce a view-generator that can be considered as a learnable augmentation for any self-supervised pre-text tasks, to generate latent viewpoint representation of a video. ViewCLR maximizes the similarities between the latent viewpoint representation with its representation from the original viewpoint, enabling the learned video encoder to generalize over unseen camera viewpoints. Experiments on cross-view benchmark datasets including NTU RGB+D dataset show that ViewCLR stands as a state-of-the-art viewpoint invariant self-supervised method.

📄 PDF Abstract BibTeX arXiv:2112.03905

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentation

Similar Papers 제목 키워드 기반

Self-Supervised Equivariant Scene Synthesis from Video

2021-02-01 · Cinjon Resnick, Or Litany, Cosmas Heiß, Hugo Larochelle 외

We propose a self-supervised framework to learn scene representations from video that are automatically delineated into background, characters, and their animations. Our method capitalizes on moving characters being equi…

Self-Supervised Video Representation Learning via Latent Time Navigation

2023-05-10 · Di Yang, Yaohui Wang, Quan Kong, Antitza Dantcheva 외

Self-supervised video representation learning aimed at maximizing similarity between different temporal segments of one video, in order to enforce feature persistence over time. This leads to loss of pertinent informatio…

Action ClassificationAction RecognitionContrastive LearningRepresentation Learning

Unseen Object Segmentation in Videos via Transferable Representations

2019-01-08 · Yi-Wen Chen, Yi-Hsuan Tsai, Chu-Ya Yang, Yen-Yu Lin 외

In order to learn object segmentation models in videos, conventional methods require a large amount of pixel-wise ground truth annotations. However, collecting such supervised data is time-consuming and labor-intensive. …

ObjectSegmentationSemantic Segmentation

Transformer-based Self-Supervised Fish Segmentation in Underwater Videos

2022-06-11 · Alzayat Saleh, Marcus Sheaves, Dean Jerry, Mostafa Rahimi Azghadi

Underwater fish segmentation to estimate fish body measurements is still largely unsolved due to the complex underwater environment. Relying on fully-supervised segmentation models requires collecting per-pixel labels, w…

Representation LearningSegmentationSelf-Supervised Learning

Multiscale Video Pretraining for Long-Term Activity Forecasting

2023-07-24 · Reuben Tan, Matthias De Lange, Michael Iuzzolino, Bryan A. Plummer 외

Long-term activity forecasting is an especially challenging research problem because it requires understanding the temporal relationships between observed actions, as well as the variability and complexity of human activ…

Action AnticipationLong Term Action Anticipation