paper-with-me

홈 › Papers

GroupViT: Semantic Segmentation Emerges from Text Supervision

2022-02-22 · CVPR 2022 1 · Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, Xiaolong Wang

Grouping and recognition are important components of visual scene understanding, e.g., for object detection and semantic segmentation. With end-to-end deep learning systems, grouping of image regions usually happens implicitly via top-down supervision from pixel-level recognition labels. Instead, in this paper, we propose to bring back the grouping mechanism into deep networks, which allows semantic segments to emerge automatically with only text supervision. We propose a hierarchical Grouping Vision Transformer (GroupViT), which goes beyond the regular grid structure representation and learns to group image regions into progressively larger arbitrary-shaped segments. We train GroupViT jointly with a text encoder on a large-scale image-text dataset via contrastive losses. With only text supervision and without any pixel-level annotations, GroupViT learns to group together semantic regions and successfully transfers to the task of semantic segmentation in a zero-shot manner, i.e., without any further fine-tuning. It achieves a zero-shot accuracy of 52.3% mIoU on the PASCAL VOC 2012 and 22.4% mIoU on PASCAL Context datasets, and performs competitively to state-of-the-art transfer-learning methods requiring greater levels of supervision. We open-source our code at https://github.com/NVlabs/GroupViT .

📄 PDF Abstract BibTeX arXiv:2202.11094

Code (6)

NVlabs/GroupViT 공식 구현 pytorch
2024-MindSpore-1/Code2/tree/main/model-1/groupvit mindspore
MindSpore-scientific-2/code-14/tree/main/groupvit mindspore
huggingface/transformers pytorch
pwc-1/Paper-9/tree/main/groupvit mindspore
yangyucheng000/University/tree/main/model-2/groupvit mindspore

Tasks

Object DetectionScene UnderstandingSemantic SegmentationTransfer LearningUnsupervised Semantic Segmentation with Language-image Pre-training

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Residual Connection 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Multi-Head Attention 설명 없음
Adam 설명 없음
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Position-Wise Feed-Forward Layer 설명 없음

Similar Papers 제목 키워드 기반

MixReorg: Cross-Modal Mixed Patch Reorganization is a Good Mask Learner for Open-World Semantic Segmentation

2023-08-09 · ICCV 2023 1 · Kaixin Cai, Pengzhen Ren, Yi Zhu, Hang Xu 외

Recently, semantic segmentation models trained with image-level text supervision have shown promising results in challenging open-world scenarios. However, these models still face difficulties in learning fine-grained se…

SegmentationSemantic SegmentationZero-Shot Semantic Segmentation

ColorFoil: Investigating Color Blindness in Large Vision and Language Models

2024-05-19 · Ahnaf Mozib Samin, M. Firoz Ahmed, Md. Mushtaq Shahriyar Rafee

With the utilization of Transformer architecture, large Vision and Language (V&L) models have shown promising performance in even zero-shot settings. Several studies, however, indicate a lack of robustness of the models …

A Semantic Space is Worth 256 Language Descriptions: Make Stronger Segmentation Models with Descriptive Properties

2023-12-21 · Junfei Xiao, Ziqi Zhou, Wenxuan Li, Shiyi Lan 외

This paper introduces ProLab, a novel approach using property-level label space for creating strong interpretable segmentation models. Instead of relying solely on category-specific annotations, ProLab uses descriptive p…

Common Sense ReasoningDescriptiveSegmentation

Enforcing View-Consistency in Class-Agnostic 3D Segmentation Fields

2024-08-19 · Corentin Dumery, Aoxiang Fan, Ren Li, Nicolas Talabot 외

Radiance Fields have become a powerful tool for modeling 3D scenes from multiple images. However, they remain difficult to segment into semantically meaningful regions. Some methods work well using 2D semantic masks, but…

Contrastive LearningObjectObject DiscoverySegmentation+1

ViewCo: Discovering Text-Supervised Segmentation Masks via Multi-View Semantic Consistency

2023-01-31 · Pengzhen Ren, Changlin Li, Hang Xu, Yi Zhu 외

Recently, great success has been made in learning visual representations from text supervision, facilitating the emergence of text-supervised semantic segmentation. However, existing works focus on pixel grouping and cro…

SegmentationSemantic Segmentation