Learning Semantics-aware Distance Map with Semantics Layering Network for Amodal Instance Segmentation
In this work, we demonstrate yet another approach to tackle the amodal segmentation problem. Specifically, we first introduce a new representation, namely a semantics-aware distance map (sem-dist map), to serve as our target for amodal segmentation instead of the commonly used masks and heatmaps. The sem-dist map is a kind of level-set representation, of which the different regions of an object are placed into different levels on the map according to their visibility. It is a natural extension of masks and heatmaps, where modal, amodal segmentation, as well as depth order information, are all well-described. Then we also introduce a novel convolutional neural network (CNN) architecture, which we refer to as semantic layering network, to estimate sem-dist maps layer by layer, from the global-level to the instance-level, for all objects in an image. Extensive experiments on the COCOA and D2SA datasets have demonstrated that our framework can predict amodal segmentation, occlusion and depth order with state-of-the-art performance.
Code (1)
Tasks
Amodal Instance SegmentationInstance SegmentationSegmentationSemantic SegmentationSimilar Papers 제목 키워드 기반
AutoLay: Benchmarking amodal layout estimation for autonomous driving
Given an image or a video captured from a monocular camera, amodal layout estimation is the task of predicting semantics and occupancy in bird's eye view. The term amodal implies we also reason about entities in the scen…
Amodal Layout EstimationAutonomous DrivingBenchmarkingEmbodied Amodal Recognition: Learning to Move to Perceive Objects
Passive visual systems typically fail to recognize objects in the amodal setting where they are heavily occluded. In contrast, humans and other embodied agents have the ability to move in the environment and actively con…
ObjectObject LocalizationSemantic SegmentationPerceiving the Invisible: Proposal-Free Amodal Panoptic Segmentation
Amodal panoptic segmentation aims to connect the perception of the world to its cognitive understanding. It entails simultaneously predicting the semantic labels of visible scene regions and the entire shape of traffic p…
Amodal Panoptic SegmentationDecoderPanoptic SegmentationEmbodied Visual Recognition
Passive visual systems typically fail to recognize objects in the amodal setting where they are heavily occluded. In contrast, humans and other embodied agents have the ability to move in the environment, and actively co…
ObjectObject LocalizationSemantic SegmentationSOFTooth: Semantics-Enhanced Order-Aware Fusion for Tooth Instance Segmentation
Three-dimensional (3D) tooth instance segmentation remains challenging due to crowded arches, ambiguous tooth-gingiva boundaries, missing teeth, and rare yet clinically important third molars. Native 3D methods relying o…
Instance Segmentation