paper-with-me

홈 › Papers

RESAnything: Attribute Prompting for Arbitrary Referring Segmentation

2025-05-03 · Ruiqi Wang, Hao Zhang

We present an open-vocabulary and zero-shot method for arbitrary referring expression segmentation (RES), targeting input expressions that are more general than what prior works were designed to handle. Specifically, our inputs encompass both object- and part-level labels as well as implicit references pointing to properties or qualities of object/part function, design, style, material, etc. Our model, coined RESAnything, leverages Chain-of-Thoughts (CoT) reasoning, where the key idea is attribute prompting. We generate detailed descriptions of object/part attributes including shape, color, and location for potential segment proposals through systematic prompting of a large language model (LLM), where the proposals are produced by a foundational image segmentation model. Our approach encourages deep reasoning about object or part attributes related to function, style, design, etc., enabling the system to handle implicit queries without any part annotations for training or fine-tuning. As the first zero-shot and LLM-based RES method, RESAnything achieves clearly superior performance among zero-shot methods on traditional RES benchmarks and significantly outperforms existing methods on challenging scenarios involving implicit queries and complex part-level relations. Finally, we contribute a new benchmark dataset to offer ~3K carefully curated RES instances to assess part-level, arbitrary RES solutions.

📄 PDF Abstract BibTeX arXiv:2505.02867

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeImage SegmentationLarge Language ModelObjectReferring ExpressionReferring Expression SegmentationSemantic Segmentation

Similar Papers 제목 키워드 기반

Visual In-Context Prompting

2023-11-22 · CVPR 2024 1 · Feng Li, Qing Jiang, Hao Zhang, Tianhe Ren 외

In-context prompting in large language models (LLMs) has become a prevalent approach to improve zero-shot capabilities, but this idea is less explored in the vision domain. Existing visual prompting methods focus on refe…

DecoderSegmentationVisual Prompting

EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

2024-06-28 · Yuxuan Zhang, Tianheng Cheng, Rui Hu, Lei Liu 외

Segment Anything Model (SAM) has attracted widespread attention for its superior interactive segmentation capabilities with visual prompts while lacking further exploration of text prompts. In this paper, we empirically …

Interactive SegmentationLanguage ModelingLanguage ModellingReferring Expression+2

Temporal Prompting Matters: Rethinking Referring Video Object Segmentation

2025-10-08 · Ci-Siang Lin, Min-Hung Chen, I-Jieh Liu, Chien-Yi Wang 외 arxiv

Referring Video Object Segmentation (RVOS) aims to segment the object referred to by the query sentence in the video. Most existing methods require end-to-end training with dense mask annotations, which could be computat…

Referring Video Object Segmentation

Attribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation

2026-07-09 · Shun Liu, Nan Xi, Yang Liu, Tianyu Luan 외 arxiv

Referring Image Segmentation (RIS) aims to segment image regions specified by natural language, enabling fine-grained and controllable visual understanding. Extending RIS to endoscopic imagery, however, presents unique c…

Image Segmentation

Learning Visual Grounding from Generative Vision and Language Model

2024-07-18 · Shijie Wang, Dahun Kim, Ali Taalimi, Chen Sun 외

Visual grounding tasks aim to localize image regions based on natural language references. In this work, we explore whether generative VLMs predominantly trained on image-text data could be leveraged to scale up the text…

AttributeLanguage ModelingLanguage ModellingObject+6