Deep Object Co-segmentation via Spatial-Semantic Network Modulation
Object co-segmentation is to segment the shared objects in multiple relevant images, which has numerous applications in computer vision. This paper presents a spatial and semantic modulated deep network framework for object co-segmentation. A backbone network is adopted to extract multi-resolution image features. With the multi-resolution features of the relevant images as input, we design a spatial modulator to learn a mask for each image. The spatial modulator captures the correlations of image feature descriptors via unsupervised learning. The learned mask can roughly localize the shared foreground object while suppressing the background. For the semantic modulator, we model it as a supervised image classification task. We propose a hierarchical second-order pooling module to transform the image features for classification use. The outputs of the two modulators manipulate the multi-resolution features by a shift-and-scale operation so that the features focus on segmenting co-object regions. The proposed model is trained end-to-end without any intricate post-processing. Extensive experiments on four image co-segmentation benchmark datasets demonstrate the superior accuracy of the proposed method compared to state-of-the-art methods.
Code (1)
Tasks
General Classificationimage-classificationImage ClassificationObjectSegmentationSimilar Papers 제목 키워드 기반
Spatial Structure Constraints for Weakly Supervised Semantic Segmentation
The image-level label has prevailed in weakly supervised semantic segmentation tasks due to its easy availability. Since image-level labels can only indicate the existence or absence of specific categories of objects, vi…
ObjectObject LocalizationSemantic SegmentationWeakly supervised Semantic Segmentation+1Spatial Frequency Modulation for Semantic Segmentation
High spatial frequency information, including fine details like textures, significantly contributes to the accuracy of semantic segmentation. However, according to the Nyquist-Shannon Sampling Theorem, high-frequency com…
Adversarial RobustnessPanoptic SegmentationSemantic SegmentationInstance SegmentationAnchorSeg: Language Grounded Query Banks for Reasoning Segmentation
Reasoning segmentation requires models to ground complex, implicit textual queries into precise pixel-level masks. Existing approaches rely on a single segmentation token $\texttt{<SEG>}$, whose hidden state implicitly e…
CM-GAN: Image Inpainting with Cascaded Modulation GAN and Object-Aware Training
Recent image inpainting methods have made great progress but often struggle to generate plausible image structures when dealing with large holes in complex images. This is partially due to the lack of effective network s…
DecoderImage InpaintingATLANTIS: A Benchmark for Semantic Segmentation of Waterbody Images
Vision-based semantic segmentation of waterbodies and nearby related objects provides important information for managing water resources and handling flooding emergency. However, the lack of large-scale labeled training …
SegmentationSemantic Segmentation