paper-with-me

홈 › Papers

Semantic Segmentation by Early Region Proxy

2022-03-26 · CVPR 2022 1 · Yifan Zhang, Bo Pang, Cewu Lu

Typical vision backbones manipulate structured features. As a compromise, semantic segmentation has long been modeled as per-point prediction on dense regular grids. In this work, we present a novel and efficient modeling that starts from interpreting the image as a tessellation of learnable regions, each of which has flexible geometrics and carries homogeneous semantics. To model region-wise context, we exploit Transformer to encode regions in a sequence-to-sequence manner by applying multi-layer self-attention on the region embeddings, which serve as proxies of specific regions. Semantic segmentation is now carried out as per-region prediction on top of the encoded region embeddings using a single linear classifier, where a decoder is no longer needed. The proposed RegProxy model discards the common Cartesian feature layout and operates purely at region level. Hence, it exhibits the most competitive performance-efficiency trade-off compared with the conventional dense prediction methods. For example, on ADE20K, the small-sized RegProxy-S/16 outperforms the best CNN model using 25% parameters and 4% computation, while the largest RegProxy-L/16 achieves 52.9mIoU which outperforms the state-of-the-art by 2.1% with fewer resources. Codes and models are available at https://github.com/YiF-Zhang/RegionProxy.

📄 PDF Abstract BibTeX arXiv:2203.14043

Code (1)

yif-zhang/regionproxy 공식 구현 pytorch

Tasks

DecoderSegmentationSemantic Segmentation

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Adam 설명 없음
Residual Connection 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Position-Wise Feed-Forward Layer 설명 없음

Similar Papers 제목 키워드 기반

Progressive Proxy Anchor Propagation for Unsupervised Semantic Segmentation

2024-07-17 · Hyun Seok Seong, WonJun Moon, SuBeen Lee, Jae-Pil Heo

The labor-intensive labeling for semantic segmentation has spurred the emergence of Unsupervised Semantic Segmentation. Recent studies utilize patch-wise contrastive learning based on features from image-level self-super…

Contrastive LearningSegmentationSemantic SegmentationUnsupervised Semantic Segmentation

LIBSVX: A Supervoxel Library and Benchmark for Early Video Processing

2015-12-30 · Chenliang Xu, Jason J. Corso

Supervoxel segmentation has strong potential to be incorporated into early video analysis as superpixel segmentation has in image analysis. However, there are many plausible supervoxel methods and little understanding as…

Boundary DetectionSegmentationSuperpixels

ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation

2024-08-09 · Mengcheng Lan, Chaofeng Chen, Yiping Ke, Xinjiang Wang 외

Open-vocabulary semantic segmentation requires models to effectively integrate visual representations with open-vocabulary semantic labels. While Contrastive Language-Image Pre-training (CLIP) models shine in recognizing…

Open Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationSegmentationSemantic Segmentation+1

Feature-Proxy Transformer for Few-Shot Segmentation

2022-10-13 · Jian-Wei Zhang, Yifan Sun, Yi Yang, Wei Chen

Few-shot segmentation (FSS) aims at performing semantic segmentation on novel classes given a few annotated support samples. With a rethink of recent advances, we find that the current FSS framework has deviated far from…

DecoderFew-Shot Semantic SegmentationSegmentationSemantic Segmentation

I Can See Clearly Now : Image Restoration via De-Raining

2019-01-03 · Horia Porav, Tom Bruls, Paul Newman

We present a method for improving segmentation tasks on images affected by adherent rain drops and streaks. We introduce a novel stereo dataset recorded using a system that allows one lens to be affected by real water dr…

DenoisingImage ReconstructionImage RestorationSegmentation+1