Joint Semantic Mining for Weakly Supervised RGB-D Salient Object Detection
Training saliency detection models with weak supervisions, e.g., image-level tags or captions, is appealing as it removes the costly demand of per-pixel annotations. Despite the rapid progress of RGB-D saliency detection in fully-supervised setting, it however remains an unexplored territory when only weak supervision signals are available. This paper is set to tackle the problem of weakly-supervised RGB-D salient object detection. The key insight in this effort is the idea of maintaining per-pixel pseudo-labels with iterative refinements by reconciling the multimodal input signals in our joint semantic mining (JSM). Considering the large variations in the raw depth map and the lack of explicit pixel-level supervisions, we propose spatial semantic modeling (SSM) to capture saliency-specific depth cues from the raw depth and produce depth-refined pseudo-labels. Moreover, tags and captions are incorporated via a fill-in-the-blank training in our textual semantic modeling (TSM) to estimate the confidences of competing pseudo-labels. At test time, our model involves only a light-weight sub-network of the training pipeline, i.e., it requires only an RGB image as input, thus allowing efficient inference. Extensive evaluations demonstrate the effectiveness of our approach under the weakly-supervised setting. Importantly, our method could also be adapted to work in both fully-supervised and unsupervised paradigms. In each of these scenarios, superior performance has been attained by our approach with comparing to the state-of-the-art dedicated methods. As a by-product, a CapS dataset is constructed by augmenting existing benchmark training set with additional image tags and captions.
Code (1)
Tasks
object-detectionObject DetectionRGB-D Salient Object DetectionSaliency DetectionSalient Object DetectionSimilar Papers 제목 키워드 기반
Non-Salient Region Object Mining for Weakly Supervised Semantic Segmentation
Semantic segmentation aims to classify every pixel of an input image. Considering the difficulty of acquiring dense labels, researchers have recently been resorting to weak labels to alleviate the annotation burden of se…
ObjectSegmentationSemantic SegmentationWeakly supervised Semantic Segmentation+1Weakly-supervised Instance Segmentation via Class-agnostic Learning with Salient Images
Humans have a strong class-agnostic object segmentation ability and can outline boundaries of unknown objects precisely, which motivates us to propose a box-supervised class-agnostic object segmentation (BoxCaseg) based …
Box-supervised Instance SegmentationInstance SegmentationMulti-Task LearningObject+4Associating Inter-Image Salient Instances for Weakly Supervised Semantic Segmentation
Effectively bridging between image level keyword annotations and corresponding image pixels is one of the main challenges in weakly supervised semantic segmentation. In this paper, we use an instance-level salient object…
Clusteringgraph partitioningImage-level Supervised Instance SegmentationInstance Segmentation+6Toward Joint Thing-and-Stuff Mining for Weakly Supervised Panoptic Segmentation
Panoptic segmentation aims to partition an image to object instances and semantic content for thing and stuff categories, respectively. To date, learning weakly supervised panoptic segmentation (WSPS) with only image…
Instance SegmentationMultiple Instance LearningObjectobject-detection+6Emphasizing Semantic Consistency of Salient Posture for Speech-Driven Gesture Generation
Speech-driven gesture generation aims at synthesizing a gesture sequence synchronized with the input speech signal. Previous methods leverage neural networks to directly map a compact audio representation to the gesture …
Gesture Generation