SceneNet RGB-D: 5M Photorealistic Images of Synthetic Indoor Trajectories with Ground Truth
We introduce SceneNet RGB-D, expanding the previous work of SceneNet to enable large scale photorealistic rendering of indoor scene trajectories. It provides pixel-perfect ground truth for scene understanding problems such as semantic segmentation, instance segmentation, and object detection, and also for geometric computer vision problems such as optical flow, depth estimation, camera pose estimation, and 3D reconstruction. Random sampling permits virtually unlimited scene configurations, and here we provide a set of 5M rendered RGB-D images from over 15K trajectories in synthetic layouts with random but physically simulated object poses. Each layout also has random lighting, camera trajectories, and textures. The scale of this dataset is well suited for pre-training data-driven computer vision techniques from scratch with RGB-D inputs, which previously has been limited by relatively small labelled datasets in NYUv2 and SUN RGB-D. It also provides a basis for investigating 3D scene labelling tasks by providing perfect camera poses and depth data as proxy for a SLAM system. We host the dataset at http://robotvault.bitbucket.io/scenenet-rgbd.html
Code (1)
Tasks
3D ReconstructionCamera Pose EstimationDepth EstimationInstance Segmentationobject-detectionObject DetectionOptical Flow EstimationPose EstimationScene UnderstandingSemantic SegmentationSimilar Papers 제목 키워드 기반
SceneNet RGB-D: Can 5M Synthetic Images Beat Generic ImageNet Pre-Training on Indoor Segmentation?
We introduce SceneNet RGB-D, a dataset providing pixel-perfect ground truth for scene understanding problems such as semantic segmentation, instance segmentation, and object detection. It also provides perfect camera pos…
16kCamera Pose EstimationInstance Segmentationobject-detection+6SegmATRon: Embodied Adaptive Semantic Segmentation for Indoor Environment
This paper presents an adaptive transformer model named SegmATRon for embodied image semantic segmentation. Its distinctive feature is the adaptation of model weights during inference on several images using a hybrid mul…
SegmentationSemantic SegmentationOpenRooms: An Open Framework for Photorealistic Indoor Scene Datasets
We propose a novel framework for creating large-scale photorealistic datasets of indoor scenes, with ground truth geometry, material, lighting and semantics. Our goal is to make the dataset creation process widely ac…
FrictionInverse RenderingLighting EstimationScene UnderstandingPerspectiveNet: A Scene-consistent Image Generator for New View Synthesis in Real Indoor Environments
Given a set of a reference RGBD views of an indoor environment, and a new viewpoint, our goal is to predict the view from that location. Prior work on new-view generation has predominantly focused on significantly constr…
Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding
For many fundamental scene understanding tasks, it is difficult or impossible to obtain per-pixel ground truth labels from real images. We address this challenge by introducing Hypersim, a photorealistic synthetic datase…
Multi-Task LearningScene UnderstandingSemantic Segmentation