Weakly Supervised Few-shot Object Segmentation using Co-Attention with Visual and Semantic Embeddings
Significant progress has been made recently in developing few-shot object segmentation methods. Learning is shown to be successful in few-shot segmentation settings, using pixel-level, scribbles and bounding box supervision. This paper takes another approach, i.e., only requiring image-level label for few-shot object segmentation. We propose a novel multi-modal interaction module for few-shot object segmentation that utilizes a co-attention mechanism using both visual and word embedding. Our model using image-level labels achieves 4.8% improvement over previously proposed image-level few-shot object segmentation. It also outperforms state-of-the-art methods that use weak bounding box supervision on PASCAL-5i. Our results show that few-shot segmentation benefits from utilizing word embeddings, and that we are able to perform few-shot segmentation using stacked joint visual semantic processing with weak image-level labels. We further propose a novel setup, Temporal Object Segmentation for Few-shot Learning (TOSFL) for videos. TOSFL can be used on a variety of public video data such as Youtube-VOS, as demonstrated in both instance-level and category-level TOSFL experiments.
Code (0)
등록된 구현이 없습니다.
Tasks
Few-Shot LearningObjectOne-shot visual object segmentationSegmentationSemantic SegmentationVideo Semantic SegmentationWord EmbeddingsSimilar Papers 제목 키워드 기반
One-Shot Weakly Supervised Video Object Segmentation
Conventional few-shot object segmentation methods learn object segmentation from a few labelled support images with strongly labelled segmentation masks. Recent work has shown to perform on par with weaker levels of supe…
ObjectSegmentationSemantic SegmentationVideo Object Segmentation+2A Language-Guided Benchmark for Weakly Supervised Open Vocabulary Semantic Segmentation
Increasing attention is being diverted to data-efficient problem settings like Open Vocabulary Semantic Segmentation (OVSS) which deals with segmenting an arbitrary object that may or may not be seen during training. The…
Few-Shot Semantic SegmentationLanguage ModellingOpen Vocabulary Semantic SegmentationOpen-Vocabulary Semantic Segmentation+3Segment, Select, Correct: A Framework for Weakly-Supervised Referring Segmentation
Referring Image Segmentation (RIS) - the problem of identifying objects in images through natural language sentences - is a challenging task currently mostly solved through supervised learning. However, while collecting …
Image SegmentationSemantic SegmentationZero-Shot LearningIntegral Object Mining via Online Attention Accumulation
Object attention maps generated by image classifiers are usually used as priors for weakly-supervised segmentation approaches. However, normal image classifiers produce attention only at the most discriminative object pa…
General ClassificationObjectSegmentationSemantic Segmentation+3Decoupled Spatial Neural Attention for Weakly Supervised Semantic Segmentation
Weakly supervised semantic segmentation receives much research attention since it alleviates the need to obtain a large amount of dense pixel-wise ground-truth annotations for the training images. Compared with other for…
Image CaptioningSegmentationSemantic SegmentationWeakly supervised Semantic Segmentation+1