paper-with-me

홈 › Papers

Temporal-aware Hierarchical Mask Classification for Video Semantic Segmentation

2023-09-14 · Zhaochong An, Guolei Sun, Zongwei Wu, Hao Tang, Luc van Gool

Modern approaches have proved the huge potential of addressing semantic segmentation as a mask classification task which is widely used in instance-level segmentation. This paradigm trains models by assigning part of object queries to ground truths via conventional one-to-one matching. However, we observe that the popular video semantic segmentation (VSS) dataset has limited categories per video, meaning less than 10% of queries could be matched to receive meaningful gradient updates during VSS training. This inefficiency limits the full expressive potential of all queries.Thus, we present a novel solution THE-Mask for VSS, which introduces temporal-aware hierarchical object queries for the first time. Specifically, we propose to use a simple two-round matching mechanism to involve more queries matched with minimal cost during training while without any extra cost during inference. To support our more-to-one assignment, in terms of the matching results, we further design a hierarchical loss to train queries with their corresponding hierarchy of primary or secondary. Moreover, to effectively capture temporal information across frames, we propose a temporal aggregation decoder that fits seamlessly into the mask-classification paradigm for VSS. Utilizing temporal-sensitive multi-level queries, our method achieves state-of-the-art performance on the latest challenging VSS benchmark VSPW without bells and whistles.

📄 PDF Abstract BibTeX arXiv:2309.08020

Code (1)

zhaochongan/the-mask 공식 구현 pytorch

Tasks

ClassificationDecoderSegmentationSemantic SegmentationVideo Semantic Segmentation

Similar Papers 제목 키워드 기반

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation

2025-01-23 · Fu Rong, Meng Lan, Qian Zhang, Lefei Zhang

Referring video object segmentation (RVOS) aims to segment objects in a video according to textual descriptions, which requires the integration of multimodal information and temporal dynamics perception. The Segment Anyt…

Referring Expression SegmentationReferring Video Object SegmentationSemantic SegmentationVideo Object Segmentation+2

TVC: Tokenized Video Compression with Ultra-Low Bitrate

2025-04-22 · Lebin Zhou, Cihan Ruan, Nam Ling, Wei Wang 외

Tokenized visual representations have shown great promise in image compression, yet their extension to video remains underexplored due to the challenges posed by complex temporal dynamics and stringent bitrate constraint…

DecoderImage CompressionVideo Compression

LoVoRA: Text-guided and Mask-free Video Object Removal and Addition with Learnable Object-aware Localization

2025-12-02 · Zhihan Xiao, Lin Liu, Yixin Gao, Xiaopeng Zhang 외 arxiv

Text-guided video editing, particularly for object removal and addition, remains a challenging task due to the need for precise spatial and temporal consistency. Existing methods often rely on auxiliary masks or referenc…

Video Inpainting

Hierarchical Attention Diffusion Networks with Object Priors for Video Change Detection

2024-08-20 · Andrew Kiruluta, Eric Lundy, Andreas Lemos

We present a unified change detection pipeline that combines instance level masking, multi\-scale attention within a denoising diffusion model, and per pixel semantic classification, all refined via SSIM to match human p…

Change DetectionDenoisingSSIM

Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation

2024-12-10 · Thong Thanh Nguyen, Xiaobao Wu, Yi Bin, Cong-Duy T Nguyen 외

To equip artificial intelligence with a comprehensive understanding towards a temporal world, video and 4D panoptic scene graph generation abstracts visual data into nodes to represent entities and edges to capture tempo…

Contrastive LearningGraph GenerationPanoptic Scene Graph GenerationRelation+2