paper-with-me

홈 › Papers

Open-Vocabulary 3D Semantic Segmentation with Text-to-Image Diffusion Models

2024-07-18 · Xiaoyu Zhu, Hao Zhou, Pengfei Xing, Long Zhao, Hao Xu, Junwei Liang, Alexander Hauptmann, Ting Liu, Andrew Gallagher

In this paper, we investigate the use of diffusion models which are pre-trained on large-scale image-caption pairs for open-vocabulary 3D semantic understanding. We propose a novel method, namely Diff2Scene, which leverages frozen representations from text-image generative models, along with salient-aware and geometric-aware masks, for open-vocabulary 3D semantic segmentation and visual grounding tasks. Diff2Scene gets rid of any labeled 3D data and effectively identifies objects, appearances, materials, locations and their compositions in 3D scenes. We show that it outperforms competitive baselines and achieves significant improvements over state-of-the-art methods. In particular, Diff2Scene improves the state-of-the-art method on ScanNet200 by 12%.

📄 PDF Abstract BibTeX arXiv:2407.13642

Code (0)

등록된 구현이 없습니다.

Tasks

3D Semantic SegmentationSemantic SegmentationVisual Grounding

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

USE: Universal Segment Embeddings for Open-Vocabulary Image Segmentation

2024-06-07 · CVPR 2024 1 · Xiaoqi Wang, Wenbin He, Xiwei Xuan, Clint Sebastian 외

The open-vocabulary image segmentation task involves partitioning images into semantically meaningful segments and classifying them with flexible text-defined categories. The recent vision-based foundation models such as…

Image SegmentationSegmentationSemantic Segmentation

Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models

2023-03-08 · CVPR 2023 1 · Jiarui Xu, Sifei Liu, Arash Vahdat, Wonmin Byeon 외

We present ODISE: Open-vocabulary DIffusion-based panoptic SEgmentation, which unifies pre-trained text-image diffusion and discriminative models to perform open-vocabulary panoptic segmentation. Text-to-image diffusion …

Open Vocabulary Panoptic SegmentationOpen Vocabulary Semantic SegmentationOpen-World Instance SegmentationPanoptic Segmentation+3

Scaling Open-Vocabulary Image Segmentation with Image-Level Labels

2021-12-22 · Golnaz Ghiasi, Xiuye Gu, Yin Cui, Tsung-Yi Lin

We design an open-vocabulary image segmentation model to organize an image into meaningful regions indicated by arbitrary texts. Recent works (CLIP and ALIGN), despite attaining impressive open-vocabulary classification …

Image SegmentationSegmentationSemantic Segmentation

Auto-Vocabulary Semantic Segmentation

2023-12-07 · Osman Ülger, Maksymilian Kulicki, Yuki Asano, Martin R. Oswald

Open-ended image understanding tasks gained significant attention from the research community, particularly with the emergence of Vision-Language Models. Open-Vocabulary Segmentation (OVS) methods are capable of performi…

Language ModelingLanguage ModellingLarge Language ModelOpen Vocabulary Semantic Segmentation+2

OVOSE: Open-Vocabulary Semantic Segmentation in Event-Based Cameras

2024-08-18 · Muhammad Rameez Ur Rahman, Jhony H. Giraldo, Indro Spinelli, Stéphane Lathuilière 외

Event cameras, known for low-latency operation and superior performance in challenging lighting conditions, are suitable for sensitive computer vision tasks such as semantic segmentation in autonomous driving. However, c…

Autonomous DrivingDomain AdaptationKnowledge DistillationOpen Vocabulary Semantic Segmentation+4