Dual Prototype Attention for Unsupervised Video Object Segmentation
Unsupervised video object segmentation (VOS) aims to detect and segment the most salient object in videos. The primary techniques used in unsupervised VOS are 1) the collaboration of appearance and motion information; and 2) temporal fusion between different frames. This paper proposes two novel prototype-based attention mechanisms, inter-modality attention (IMA) and inter-frame attention (IFA), to incorporate these techniques via dense propagation across different modalities and frames. IMA densely integrates context information from different modalities based on a mutual refinement. IFA injects global context of a video to the query frame, enabling a full utilization of useful properties from multiple frames. Experimental results on public benchmark datasets demonstrate that our proposed approach outperforms all existing methods by a substantial margin. The proposed two components are also thoroughly validated via ablative study.
Code (1)
Tasks
ObjectSemantic SegmentationUnsupervised Video Object SegmentationVideo Object SegmentationVideo Semantic SegmentationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Unsupervised Video Object Segmentation via Prototype Memory Network
Unsupervised video object segmentation aims to segment a target object in the video without a ground truth mask in the initial frame. This challenging task requires extracting features for the most salient common objects…
ObjectOptical Flow EstimationSelf-LearningSemantic Segmentation+3Holistic Prototype Attention Network for Few-Shot VOS
Few-shot video object segmentation (FSVOS) aims to segment dynamic objects of unseen classes by resorting to a small set of support images that contain pixel-level object annotations. Existing methods have demonstrated t…
Graph AttentionSemantic SegmentationVideo Object SegmentationVideo Semantic SegmentationEfficient Unsupervised Video Object Segmentation Network Based on Motion Guidance
Due to the problem of performance constraints of unsupervised video object detection, its large-scale application is limited. In response to this pain point, we propose another excellent method to solve this problematic …
object-detectionObject DetectionOptical Flow EstimationSemantic Segmentation+4VONet: Unsupervised Video Object Learning With Parallel U-Net Attention and Object-wise Sequential VAE
Unsupervised video object learning seeks to decompose video scenes into structural object representations without any supervision from depth, optical flow, or segmentation. We present VONet, an innovative approach that i…
DecoderObjectOptical Flow EstimationUnsupervised Open-Vocabulary Object Localization in Videos
In this paper, we show that recent advances in video representation learning and pre-trained vision-language models allow for substantial improvements in self-supervised video object localization. We propose a method tha…
ObjectObject LocalizationRepresentation Learning