Semantic-Aware Generation for Self-Supervised Visual Representation Learning
In this paper, we propose a self-supervised visual representation learning approach which involves both generative and discriminative proxies, where we focus on the former part by requiring the target network to recover the original image based on the mid-level features. Different from prior work that mostly focuses on pixel-level similarity between the original and generated images, we advocate for Semantic-aware Generation (SaGe) to facilitate richer semantics rather than details to be preserved in the generated image. The core idea of implementing SaGe is to use an evaluator, a deep network that is pre-trained without labels, for extracting semantic-aware features. SaGe complements the target network with view-specific features and thus alleviates the semantic degradation brought by intensive data augmentations. We execute SaGe on ImageNet-1K and evaluate the pre-trained models on five downstream tasks including nearest neighbor test, linear classification, and fine-scaled image recognition, demonstrating its ability to learn stronger visual representations.
Code (1)
Tasks
Representation LearningSemantic SegmentationSimilar Papers 제목 키워드 기반
Semantic-Aware Fine-Grained Correspondence
Establishing visual correspondence across images is a challenging and essential task. Recently, an influx of self-supervised methods have been proposed to better learn representations for visual correspondence. However, …
Pose TrackingSelf-Supervised LearningSemantic correspondenceSemantic Segmentation+2Understanding Self-Supervised Features for Learning Unsupervised Instance Segmentation
Self-supervised learning (SSL) can be used to solve complex visual tasks without human labels. Self-supervised representations encode useful semantic information about images, and as a result, they have already been used…
Instance SegmentationSegmentationSelf-Supervised LearningSemantic Segmentation+3PASS: Patch-Aware Self-Supervision for Vision Transformer
Recent self-supervised representation learning methods have shown impressive results in learning visual representations from unlabeled images. This paper aims to improve their performance further by utilizing the archite…
object-detectionObject DetectionRepresentation LearningSelf-Supervised Learning+1VLPD: Context-Aware Pedestrian Detection via Vision-Language Semantic Self-Supervision
Detecting pedestrians accurately in urban scenes is significant for realistic applications like autonomous driving or video surveillance. However, confusing human-like objects often lead to wrong detections, and small sc…
Autonomous DrivingPedestrian DetectionGaussian Masked Autoencoders
This paper explores Masked Autoencoders (MAE) with Gaussian Splatting. While reconstructive self-supervised learning frameworks such as MAE learns good semantic abstractions, it is not trained for explicit spatial awaren…
Edge DetectionRepresentation LearningSelf-Supervised LearningZero-Shot Learning