paper-with-me

Papers

VSPW: A Large-scale Dataset for Video Scene Parsing in the Wild

2021-06-19 · CVPR 2021 1 · Jiaxu Miao, Yunchao Wei, Yu Wu, Chen Liang, Guangrui Li, Yi Yang

In this paper, we present a new dataset with the target of advancing the scene parsing task from images to videos. Our dataset aims to perform Video Scene Parsing in the Wild (VSPW), which covers a wide range of real-world scenarios and categories. To be specific, our VSPW is featured from the following aspects: 1) Well-trimmed long-temporal clips. Each video contains a complete shot, lasting around 5 seconds on average. 2) Dense annotation. The pixel-level annotations are provided at a high frame rate of 15 f/s. 3) High resolution. Over 96% of the captured videos are with high spatial resolutions from 720P to 4K. We totally annotate 3,337 videos, including 239,934 frames from 124 categories. To the best of our knowledge, our VSPW is the first attempt to tackle the challenging video scene parsing task in the wild by considering diverse scenarios. Based on VSPW, we design a generic Temporal Context Blending (TCB) network, which can effectively harness long-range contextual information from the past frames to help segment the current one. Extensive experiments show that our TCB network improves both the segmentation performance and temporal stability comparing with image-/video-based state-of-the-art methods. We hope that the scale, diversity, long-temporal, and high frame rate of our VSPW can significantly advance the research of video scene parsing and beyond.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

4kScene Parsing

Similar Papers 제목 키워드 기반

Semantic Segmentation on VSPW Dataset through Contrastive Loss and Multi-dataset Training Approach

2023-06-06 · Min Yan, Qianxiong Ning, Qian Wang

Video scene parsing incorporates temporal information, which can enhance the consistency and accuracy of predictions compared to image scene parsing. The added temporal dimension enables a more comprehensive understandin…

Scene ParsingSemantic SegmentationVideo Semantic Segmentation

TBN-ViT: Temporal Bilateral Network with Vision Transformer for Video Scene Parsing

2021-12-02 · Bo Yan, Leilei Cao, Hongbin Wang

Video scene parsing in the wild with diverse scenarios is a challenging and great significance task, especially with the rapid development of automatic driving technique. The dataset Video Scene Parsing in the Wild(VSPW)…

Scene Parsing

Exploiting Spatial-Temporal Semantic Consistency for Video Scene Parsing

2021-09-06 · Xingjian He, Weining Wang, Zhiyong Xu, Hao Wang 외

Compared with image scene parsing, video scene parsing introduces temporal information, which can effectively improve the consistency and accuracy of prediction. In this paper, we propose a Spatial-Temporal Semantic Cons…

Scene Parsing

1st Place Winner of the 2024 Pixel-level Video Understanding in the Wild (CVPR'24 PVUW) Challenge in Video Panoptic Segmentation and Best Long Video Consistency of Video Semantic Segmentation

2024-06-08 · Qingfeng Liu, Mostafa El-Khamy, Kee-Bong Song

The third Pixel-level Video Understanding in the Wild (PVUW CVPR 2024) challenge aims to advance the state of art in video understanding through benchmarking Video Panoptic Segmentation (VPS) and Video Semantic Segmentat…

BenchmarkingInstance SegmentationPanoptic SegmentationScene Parsing+6

Semantic Segmentation on VSPW Dataset through Aggregation of Transformer Models

2021-09-03 · Zixuan Chen, Junhong Zou, Xiaotao Wang

Semantic segmentation is an important task in computer vision, from which some important usage scenarios are derived, such as autonomous driving, scene parsing, etc. Due to the emphasis on the task of video semantic segm…

Autonomous DrivingScene ParsingSegmentationSemantic Segmentation+1