paper-with-me

홈 › Papers

DisCLIP: Open-Vocabulary Referring Expression Generation

2023-05-30 · Lior Bracha, Eitan Shaar, Aviv Shamsian, Ethan Fetaya, Gal Chechik

Referring Expressions Generation (REG) aims to produce textual descriptions that unambiguously identifies specific objects within a visual scene. Traditionally, this has been achieved through supervised learning methods, which perform well on specific data distributions but often struggle to generalize to new images and concepts. To address this issue, we present a novel approach for REG, named DisCLIP, short for discriminative CLIP. We build on CLIP, a large-scale visual-semantic model, to guide an LLM to generate a contextual description of a target concept in an image while avoiding other distracting concepts. Notably, this optimization happens at inference time and does not require additional training or tuning of learned parameters. We measure the quality of the generated text by evaluating the capability of a receiver model to accurately identify the described object within the scene. To achieve this, we use a frozen zero-shot comprehension module as a critique of our generated referring expressions. We evaluate DisCLIP on multiple referring expression benchmarks through human evaluation and show that it significantly outperforms previous methods on out-of-domain datasets. Our results highlight the potential of using pre-trained visual-semantic models for generating high-quality contextual descriptions.

📄 PDF Abstract BibTeX arXiv:2305.19108

Code (0)

등록된 구현이 없습니다.

Tasks

Referring ExpressionReferring expression generation

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding

2026-01-14 · Sheng-Yu Huang, Jaesung Choe, Yu-Chiang Frank Wang, Cheng Sun arxiv

We propose OpenVoxel, a training-free algorithm for grouping and captioning sparse voxels for the open-vocabulary 3D scene understanding tasks. Given the sparse voxel rasterization (SVR) model obtained from multi-view im…

Referring Expression SegmentationScene Understanding

Open-Vocabulary and Referring Segmentation for 3D Gaussians Using 2D Detectors

2026-06-29 · Jameel Hassan, Yasiru Ranasinghe, Vishal Patel arxiv

3D Gaussian Splatting (3DGS) has emerged at the forefront of 3D scene reconstruction. Extending 3DGS with language-driven, open-vocabulary understanding has gained significant attention for real-world applications such a…

Referring ExpressionSpatial Reasoning

WeDetect: Fast Open-Vocabulary Object Detection as Retrieval

2025-12-13 · Shenghao Fu, Yukun Su, Fengyun Rao, Jing Lyu 외 arxiv

Open-vocabulary object detection aims to detect arbitrary classes via text prompts. Methods without cross-modal fusion layers (non-fusion) offer faster inference by treating recognition as a retrieval problem, \ie, match…

Referring ExpressionObject Detection

Beyond Similarity Matching: Structured Reasoning for Open-Vocabulary Referring Segmentation in 3DGS

2026-08-17 · Yizhao Wang, Xinfa Wang, Jingbo Wang, Jingbo Wang 외 arxiv

Open-vocabulary referring segmentation in 3D Gaussian Splatting (3DGS) requires a neural model to select Gaussian primitives according to free-form language expressions. Existing 3DGS-based methods usually rely on global…

On The Feasibility of Open Domain Referring Expression Generation Using Large Scale Folksonomies

2012-06-01 · NAACL 2012 6 · Fabi{\'a}n Pacheco, Pablo Duboue, Mart{\'\i}n Dom{\'\i}nguez
Document SummarizationMulti-Document SummarizationReferring ExpressionReferring expression generation+1