Scene Parsing with Object Instances and Occlusion Ordering
This work proposes a method to interpret a scene by assigning a semantic label at every pixel and inferring the spatial extent of individual object instances together with their occlusion relationships. Starting with an initial pixel labeling and a set of candidate object masks for a given test image, we select a subset of objects that explain the image well and have valid overlap relationships and occlusion ordering. This is done by minimizing an integer quadratic program either using a greedy method or a standard solver. Then we alternate between using the object predictions to refine the pixel labels and vice versa. The proposed system obtains promising results on two challenging subsets of the LabelMe and SUN datasets, the largest of which contains 45,676 images and 232 classes.
Code (0)
등록된 구현이 없습니다.
Tasks
ObjectScene ParsingvalidSimilar Papers 제목 키워드 기반
Instance-wise Occlusion and Depth Orders in Natural Scenes
In this paper, we introduce a new dataset, named InstaOrder, that can be used to understand the geometrical relationships of instances in an image. The dataset consists of 2.9M annotations of geometric orderings for clas…
Depth EstimationDepth PredictionPredictionScene UnderstandingBBBD: Bounding Box Based Detector for Occlusion Detection and Order Recovery
Occlusion handling is one of the challenges of object detection and segmentation, and scene understanding. Because objects appear differently when they are occluded in varying degree, angle, and locations. Therefore, det…
object-detectionObject DetectionOcclusion HandlingScene Understanding+1Self-Supervised Scene De-occlusion
Natural scene understanding is a challenging task, particularly when encountering images of multiple objects that are partially occluded. This obstacle is given rise by varying object ordering and positioning. Existing s…
Image ManipulationScene UnderstandingRevisiting Depth Layers from Occlusions
In this work, we consider images of a scene with a moving object captured by a static camera. As the object (human or otherwise) moves about the scene, it reveals pairwise depth-ordering or occlusion cues. The goal of th…
ObjectOcclusionFormer: Arranging Z-Order for Layout-Grounded Image Generation
Recent layout-to-image models have achieved remarkable progress in spatial controllability. However, they still struggle with inter-object occlusion. When bounding boxes overlap, most existing methods lack explicit occlu…
Image Generation