RendBEV: Semantic Novel View Synthesis for Self-Supervised Bird's Eye View Segmentation
Bird's Eye View (BEV) semantic maps have recently garnered a lot of attention as a useful representation of the environment to tackle assisted and autonomous driving tasks. However, most of the existing work focuses on the fully supervised setting, training networks on large annotated datasets. In this work, we present RendBEV, a new method for the self-supervised training of BEV semantic segmentation networks, leveraging differentiable volumetric rendering to receive supervision from semantic perspective views computed by a 2D semantic segmentation model. Our method enables zero-shot BEV semantic segmentation, and already delivers competitive results in this challenging setting. When used as pretraining to then fine-tune on labeled BEV ground-truth, our method significantly boosts performance in low-annotation regimes, and sets a new state of the art when fine-tuning on all available labels.
Code (0)
등록된 구현이 없습니다.
Tasks
2D Semantic SegmentationAutonomous DrivingNovel View SynthesisSegmentationSemantic SegmentationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Self-supervised Multi-view Stereo via Effective Co-Segmentation and Data-Augmentation
Recent studies have witnessed that self-supervised methods based on view synthesis obtain clear progress on multi-view stereo (MVS). However, existing methods rely on the assumption that the corresponding points among di…
Data AugmentationLearning Dense Object Descriptors from Multiple Views for Low-shot Category Generalization
A hallmark of the deep learning era for computer vision is the successful use of large-scale labeled datasets to train feature representations for tasks ranging from object recognition and semantic segmentation to optica…
Novel View SynthesisObjectObject RecognitionOptical Flow Estimation+2SODA: Bottleneck Diffusion Models for Representation Learning
We introduce SODA, a self-supervised diffusion model, designed for representation learning. The model incorporates an image encoder, which distills a source view into a compact representation, that, in turn, guides the g…
DecoderDenoisingImage GenerationLinear-Probe Classification+2SelfOcc: Self-Supervised Vision-Based 3D Occupancy Prediction
3D occupancy prediction is an important task for the robustness of vision-centric autonomous driving, which aims to predict whether each point is occupied in the surrounding 3D space. Existing methods usually require 3D …
Autonomous DrivingDepth EstimationMonocular Depth EstimationPredictionSelf-supervised Light Field View Synthesis Using Cycle Consistency
High angular resolution is advantageous for practical applications of light fields. In order to enhance the angular resolution of light fields, view synthesis methods can be utilized to generate dense intermediate views …