paper-with-me

Papers

Attribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation

2026-07-09 · Shun Liu, Nan Xi, Yang Liu, Tianyu Luan, Xuan Gong, David Doermann arxiv

Referring Image Segmentation (RIS) aims to segment image regions specified by natural language, enabling fine-grained and controllable visual understanding. Extending RIS to endoscopic imagery, however, presents unique challenges, including scarce high-quality annotations and complex, domain-specific image-text relationships. Although recent vision-language models demonstrate strong cross-domain alignment, they often fail to capture fine-grained textual cues in endoscopic settings, resulting in suboptimal performance and limited generalization. To address these challenges, we introduce ReferEndoscopy, a large-scale benchmark for RIS in the endoscopy field. Building on this dataset, we propose the Attribute Retrieval-based Endoscopic-RIS (AR-ERIS) framework for open-vocabulary endoscopic compositional referring segmentation. AR-ERIS leverages attribute retrieval for open-vocabulary endoscopic compositional referring segmentation and is pretrained on the curated ReferEndoscopy dataset, achieving state-of-the-art performance with strong generalization across both simulated and real-world endoscopic data. The dataset and code will be publicly released upon completion of the review process.

📄 PDF Abstract BibTeX arXiv:2607.08397

Code (0)

등록된 구현이 없습니다.

Tasks

Image Segmentation

Similar Papers 제목 키워드 기반

SANEval: Open-Vocabulary Compositional Benchmarks with Failure-mode Diagnosis

2026-01-30 · Rishav Pramanik, Ian E. Nielsen, Jeff Smith, Saurav Pandit 외 arxiv

The rapid progress of text-to-image (T2I) models has unlocked unprecedented creative potential, yet their ability to faithfully render complex prompts involving multiple objects, attributes, and spatial relationships rem…

Compositional Caching for Training-free Open-vocabulary Attribute Detection

2025-01-01 · CVPR 2025 1 · Marco Garosi, Alessandro Conti, Gaowen Liu, Elisa Ricci 외

Attribute detection is crucial for many computer vision tasks, as it enables systems to describe properties such as color, texture, and material. Current approaches often rely on labor-intensive annotation processes …

AttributeOpen Vocabulary Attribute Detection

Structure-aware Prompt Adaptation from Seen to Unseen for Open-Vocabulary Compositional Zero-Shot Learning

2026-03-04 · Yihang Duan, Jiong Wang, Pengpeng Zeng, Ji Zhang 외 arxiv

The goal of Open-Vocabulary Compositional Zero-Shot Learning (OV-CZSL) is to recognize attribute-object compositions in the open-vocabulary setting, where compositions of both seen and unseen attributes and objects are e…

Compositional Zero-Shot Learning

Decomposed Vision-Language Alignment for Fine-Grained Open-Vocabulary Segmentation

2026-05-15 · Chenhao Wang, Yingrui Ji, Yu Meng, Yao Zhu arxiv

Open-vocabulary segmentation models often struggle to generalize to unseen combinations of object categories and attributes, because fine-grained descriptions are typically encoded as holistic sentences that entangle mul…

Beyond Seen Primitive Concepts and Attribute-Object Compositional Learning

2024-01-01 · CVPR 2024 1 · Nirat Saini, Khoi Pham, Abhinav Shrivastava

Learning from seen attribute-object pairs to generalize to unseen compositions has been studied extensively in Compositional Zero-Shot Learning (CZSL). However CZSL setup is still limited to seen attributes and objec…

AttributeCompositional Zero-Shot LearningZero-Shot Learning