paper-with-me

홈 › Papers

Generalizable Semantic Vision Query Generation for Zero-shot Panoptic and Semantic Segmentation

2024-02-21 · Jialei Chen, Daisuke Deguchi, Chenkai Zhang, Hiroshi Murase

Zero-shot Panoptic Segmentation (ZPS) aims to recognize foreground instances and background stuff without images containing unseen categories in training. Due to the visual data sparsity and the difficulty of generalizing from seen to unseen categories, this task remains challenging. To better generalize to unseen classes, we propose Conditional tOken aligNment and Cycle trAnsiTion (CONCAT), to produce generalizable semantic vision queries. First, a feature extractor is trained by CON to link the vision and semantics for providing target queries. Formally, CON is proposed to align the semantic queries with the CLIP visual CLS token extracted from complete and masked images. To address the lack of unseen categories, a generator is required. However, one of the gaps in synthesizing pseudo vision queries, ie, vision queries for unseen categories, is describing fine-grained visual details through semantic embeddings. Therefore, we approach CAT to train the generator in semantic-vision and vision-semantic manners. In semantic-vision, visual query contrast is proposed to model the high granularity of vision by pulling the pseudo vision queries with the corresponding targets containing segments while pushing those without segments away. To ensure the generated queries retain semantic information, in vision-semantic, the pseudo vision queries are mapped back to semantic and supervised by real semantic embeddings. Experiments on ZPS achieve a 5.2% hPQ increase surpassing SOTA. We also examine inductive ZPS and open-vocabulary semantic segmentation and obtain comparative results while being 2 times faster in testing.

📄 PDF Abstract BibTeX arXiv:2402.13697

Code (0)

등록된 구현이 없습니다.

Tasks

Open Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationPanoptic SegmentationSemantic Segmentation

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

How to Train Your DRAGON: Diverse Augmentation Towards Generalizable Dense Retrieval

2023-02-15 · Sheng-Chieh Lin, Akari Asai, Minghan Li, Barlas Oguz 외

Various techniques have been developed in recent years to improve dense retrieval (DR), such as unsupervised contrastive learning and pseudo-query generation. Existing DRs, however, often suffer from effectiveness tradeo…

Contrastive LearningData AugmentationPassage RetrievalRetrieval+1

Generalizable Object Keypoint Localization from Generative Priors

2025-01-01 · CVPR 2025 1 · Dongkai Wang, Jiang Duan, Liangjian Wen, Shiyu Xuan 외

Generalizable object keypoint localization is a fundamental computer vision task in understanding the object structure. It is challenging for existing keypoint localization methods because their limited training data…

Cross-Domain Few-ShotImage GenerationObject

g3D-LF: Generalizable 3D-Language Feature Fields for Embodied Tasks

2024-11-26 · CVPR 2025 1 · Zihan Wang, Gim Hee Lee

We introduce Generalizable 3D-Language Feature Fields (g3D-LF), a 3D representation model pre-trained on large-scale 3D-language dataset for embodied tasks. Our g3D-LF processes posed RGB-D images from agents to encode f…

Contrastive LearningQuestion AnsweringVision and Language Navigation

FeatureNeRF: Learning Generalizable NeRFs by Distilling Foundation Models

2023-03-22 · ICCV 2023 1 · Jianglong Ye, Naiyan Wang, Xiaolong Wang

Recent works on generalizable NeRFs have shown promising results on novel view synthesis from single or few images. However, such models have rarely been applied on other downstream tasks beyond synthesis such as semanti…

NeRFNeural RenderingNovel View Synthesis

Q-Align: Alleviating Attention Leakage in Zero-Shot Appearance Transfer via Query-Query Alignment

2025-08-27 · Namu Kim, Wonbin Kweon, Minsoo Kim, Hwanjo Yu arxiv

We observe that zero-shot appearance transfer with large-scale image generation models faces a significant challenge: Attention Leakage. This challenge arises when the semantic mapping between two images is captured by t…

Image Generation