Memory Based Video Scene Parsing
Video scene parsing is a long-standing challenging task in computer vision, aiming to assign pre-defined semantic labels to pixels of all frames in a given video. Compared with image semantic segmentation, this task pays more attention on studying how to adopt the temporal information to obtain higher predictive accuracy. In this report, we introduce our solution for the 1st Video Scene Parsing in the Wild Challenge, which achieves a mIoU of 57.44 and obtained the 2nd place (our team name is CharlesBLWX).
Code (0)
등록된 구현이 없습니다.
Tasks
Scene ParsingSemantic SegmentationSimilar Papers 제목 키워드 기반
Video Scene Parsing with Predictive Feature Learning
In this work, we address the challenging video scene parsing problem by developing effective representation learning methods given limited parsing annotations. In particular, we contribute two novel methods that constitu…
Representation LearningScene ParsingVSPW: A Large-scale Dataset for Video Scene Parsing in the Wild
In this paper, we present a new dataset with the target of advancing the scene parsing task from images to videos. Our dataset aims to perform Video Scene Parsing in the Wild (VSPW), which covers a wide range of real…
4kScene ParsingSemantic Segmentation on VSPW Dataset through Aggregation of Transformer Models
Semantic segmentation is an important task in computer vision, from which some important usage scenarios are derived, such as autonomous driving, scene parsing, etc. Due to the emphasis on the task of video semantic segm…
Autonomous DrivingScene ParsingSegmentationSemantic Segmentation+1Discourse Parsing in Videos: A Multi-modal Appraoch
Text-level discourse parsing aims to unmask how two sentences in the text are related to each other. We propose the task of Visual Discourse Parsing, which requires understanding discourse relations among scenes in a vid…
Discourse ParsingVisual DialogVisual StorytellingSemi-supervised Video Semantic Segmentation Using Unreliable Pseudo Labels for PVUW2024
Pixel-level Scene Understanding is one of the fundamental problems in computer vision, which aims at recognizing object classes, masks and semantics of each pixel in the given image. Compared with image scene parsing, vi…
Scene ParsingScene UnderstandingSemantic SegmentationVideo Semantic Segmentation