A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects
Video Scene Parsing (VSP) has emerged as a cornerstone in computer vision, facilitating the simultaneous segmentation, recognition, and tracking of diverse visual entities in dynamic scenes. In this survey, we present a holistic review of recent advances in VSP, covering a wide array of vision tasks, including Video Semantic Segmentation (VSS), Video Instance Segmentation (VIS), Video Panoptic Segmentation (VPS), as well as Video Tracking and Segmentation (VTS), and Open-Vocabulary Video Segmentation (OVVS). We systematically analyze the evolution from traditional hand-crafted features to modern deep learning paradigms -- spanning from fully convolutional networks to the latest transformer-based architectures -- and assess their effectiveness in capturing both local and global temporal contexts. Furthermore, our review critically discusses the technical challenges, ranging from maintaining temporal consistency to handling complex scene dynamics, and offers a comprehensive comparative study of datasets and evaluation metrics that have shaped current benchmarking standards. By distilling the key contributions and shortcomings of state-of-the-art methodologies, this survey highlights emerging trends and prospective research directions that promise to further elevate the robustness and adaptability of VSP in real-world applications.
Code (0)
등록된 구현이 없습니다.
Tasks
BenchmarkingInstance SegmentationOpen-Vocabulary Video SegmentationPanoptic SegmentationScene ParsingSegmentationSemantic SegmentationSurveyVideo Instance SegmentationVideo Panoptic SegmentationVideo SegmentationVideo Semantic SegmentationSimilar Papers 제목 키워드 기반
About Time: Advances, Challenges, and Outlooks of Action Understanding
We have witnessed impressive advances in video action understanding. Increased dataset sizes, variability, and computation availability have enabled leaps in performance and task diversification. Current systems can prov…
Action UnderstandingSurveyDeep Learning Technique for Human Parsing: A Survey and Outlook
Human parsing aims to partition humans in image or video into multiple pixel-level semantic parts. In the last decade, it has gained significantly increased interest in the computer vision community and has been utilized…
Deep LearningHuman ParsingSurveyMultimodal Referring Segmentation: A Survey
Multimodal referring segmentation aims to segment target objects in visual scenes, such as images, videos, and 3D scenes, based on referring expressions in text or audio format. This task plays a crucial role in practica…
Referring ExpressionVideo Scene Parsing with Predictive Feature Learning
In this work, we address the challenging video scene parsing problem by developing effective representation learning methods given limited parsing annotations. In particular, we contribute two novel methods that constitu…
Representation LearningScene ParsingRecent Advances in Video Question Answering: A Review of Datasets and Methods
Video Question Answering (VQA) is a recent emerging challenging task in the field of Computer Vision. Several visual information retrieval techniques like Video Captioning/Description and Video-guided Machine Translation…
Information RetrievalMachine TranslationQuestion AnsweringRetrieval+6