Scene Parsing Through ADE20K Dataset
Scene parsing, or recognizing and segmenting objects and stuff in an image, is one of the key problems in computer vision. Despite the community's efforts in data collection, there are still few image datasets covering a wide range of scenes and object categories with dense and detailed annotations for scene parsing. In this paper, we introduce and analyze the ADE20K dataset, spanning diverse annotations of scenes, objects, parts of objects, and in some cases even parts of parts. A scene parsing benchmark is built upon the ADE20K with 150 object and stuff classes included. Several segmentation baseline models are evaluated on the benchmark. A novel network design called Cascade Segmentation Module is proposed to parse a scene into stuff, objects, and object parts in a cascade and improve over the baselines. We further show that the trained scene parsing networks can lead to applications such as image content removal and scene synthesis(Dataset and pretrained models are available at http://groups.csail.mit.edu/vision/datasets/ADE20K/).
Code (0)
등록된 구현이 없습니다.
Tasks
ObjectScene ParsingSegmentationSimilar Papers 제목 키워드 기반
Pyramid Scene Parsing Network
Scene parsing is challenging for unrestricted open vocabulary and diverse scenes. In this paper, we exploit the capability of global context information by different-region-based context aggregation through our pyramid p…
Dichotomous Image SegmentationImage ClassificationLesion SegmentationReal-Time Semantic Segmentation+4Semantic Understanding of Scenes through the ADE20K Dataset
Scene parsing, or recognizing and segmenting objects and stuff in an image, is one of the key problems in computer vision. Despite the community's efforts in data collection, there are still few image datasets covering a…
Scene ParsingSegmentationSemantic SegmentationTraffic Scene Parsing through the TSP6K Dataset
Traffic scene perception in computer vision is a critically important task to achieve intelligent cities. To date, most existing datasets focus on autonomous driving scenes. We observe that the models trained on those dr…
Autonomous DrivingDecoderDomain AdaptationInstance Segmentation+3FoveaNet: Perspective-aware Urban Scene Parsing
Parsing urban scene images benefits many applications, especially self-driving. Most of the current solutions employ generic image parsing models that treat all scales and locations in the images equally and do not consi…
Scene ParsingSemantic Segmentation on VSPW Dataset through Aggregation of Transformer Models
Semantic segmentation is an important task in computer vision, from which some important usage scenarios are derived, such as autonomous driving, scene parsing, etc. Due to the emphasis on the task of video semantic segm…
Autonomous DrivingScene ParsingSegmentationSemantic Segmentation+1