paper-with-me

홈 › Papers

Text and Click inputs for unambiguous open vocabulary instance segmentation

2023-11-24 · Nikolai Warner, Meera Hahn, Jonathan Huang, Irfan Essa, Vighnesh Birodkar

Segmentation localizes objects in an image on a fine-grained per-pixel scale. Segmentation benefits by humans-in-the-loop to provide additional input of objects to segment using a combination of foreground or background clicks. Tasks include photoediting or novel dataset annotation, where human annotators leverage an existing segmentation model instead of drawing raw pixel level annotations. We propose a new segmentation process, Text + Click segmentation, where a model takes as input an image, a text phrase describing a class to segment, and a single foreground click specifying the instance to segment. Compared to previous approaches, we leverage open-vocabulary image-text models to support a wide-range of text prompts. Conditioning segmentations on text prompts improves the accuracy of segmentations on novel or unseen classes. We demonstrate that the combination of a single user-specified foreground click and a text prompt allows a model to better disambiguate overlapping or co-occurring semantic categories, such as "tie", "suit", and "person". We study these results across common segmentation datasets such as refCOCO, COCO, VOC, and OpenImages. Source code available here.

📄 PDF Abstract BibTeX arXiv:2311.14822

Code (1)

nikolaiwarner7/text-and-click-for-open-vocabulary-segmentation 공식 구현

Tasks

Instance SegmentationSegmentationSemantic Segmentation

Similar Papers 제목 키워드 기반

Open 3D World in Autonomous Driving

2024-08-20 · Xinlong Cheng, Lei LI

The capability for open vocabulary perception represents a significant advancement in autonomous driving systems, facilitating the comprehension and interpretation of a wide array of textual inputs in real-time. Despite …

Autonomous DrivingAutonomous Navigation

OpenGaussian: Towards Point-Level 3D Gaussian-based Open Vocabulary Understanding

2024-06-04 · Yanmin Wu, Jiarui Meng, Haijie Li, Chenming Wu 외

This paper introduces OpenGaussian, a method based on 3D Gaussian Splatting (3DGS) capable of 3D point-level open vocabulary understanding. Our primary motivation stems from observing that existing 3DGS-based open vocabu…

3DGSObject

Enabling Training-Free Text-Based Remote Sensing Segmentation

2026-02-19 · Jose Sosa, Danila Rukhovich, Anis Kacem, Djamila Aouada arxiv

Recent advances in Vision Language Models (VLMs) and Vision Foundation Models (VFMs) have opened new opportunities for zero-shot text-guided segmentation of remote sensing imagery. However, most existing approaches still…

Semantic Segmentation

Just Add Geometry: Gradient-Free Open-Vocabulary 3D Detection Without Human-in-the-Loop

2025-07-06 · Atharv Goel, Mehar Khurana arxiv

Modern 3D object detection datasets are constrained by narrow class taxonomies and costly manual annotations, limiting their ability to scale to open-world settings. In contrast, 2D vision-language models trained on web-…

3D Object Detection

Follow Anything: Open-set detection, tracking, and following in real-time

2023-08-10 · Alaa Maalouf, Ninad Jadhav, Krishna Murthy Jatavallabhula, Makram Chahine 외

Tracking and following objects of interest is critical to several robotics use cases, ranging from industrial automation to logistics and warehousing, to healthcare and security. In this paper, we present a robotic syste…