paper-with-me

Papers

Text-guided Zero-Shot Object Localization

2024-11-18 · Jingjing Wang, Xinglin Piao, Zongzhi Gao, Bo Li, Yong Zhang, BaoCai Yin

Object localization is a hot issue in computer vision area, which aims to identify and determine the precise location of specific objects from image or video. Most existing object localization methods heavily rely on extensive labeled data, which are costly to annotate and constrain their applicability. Therefore, we propose a new Zero-Shot Object Localization (ZSOL) framework for addressing the aforementioned challenges. In the proposed framework, we introduce the Contrastive Language Image Pre-training (CLIP) module which could integrate visual and linguistic information effectively. Furthermore, we design a Text Self-Similarity Matching (TSSM) module, which could improve the localization accuracy by enhancing the representation of text features extracted by CLIP module. Hence, the proposed framework can be guided by prompt words to identify and locate specific objects in an image in the absence of labeled samples. The results of extensive experiments demonstrate that the proposed method could improve the localization performance significantly and establishes an effective benchmark for further research.

📄 PDF Abstract BibTeX arXiv:2411.11357

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectObject Localization

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Semantic-Guided Multi-Attention Localization for Zero-Shot Learning

2019-03-01 · NeurIPS 2019 12 · Yizhe Zhu, Jianwen Xie, Zhiqiang Tang, Xi Peng 외

Zero-shot learning extends the conventional object classification to the unseen class recognition by introducing semantic representations of classes. Existing approaches predominantly focus on learning the proper mapping…

TripletZero-Shot Learning

LAGO: Language-Guided Adaptive Object-Region Focus for Zero-Shot Visual-Text Alignment

2026-05-04 · Junyi Hu, Qiji Zhou, Lei Zhang, Yue Zhang arxiv

Zero-shot recognition aims to classify an image by selecting the most compatible label description from a set of candidate classes without any task-specific supervision. In fine-grained settings, however, the relevant ev…

DiffuSAM: Diffusion Guided Zero-Shot Object Grounding for Remote Sensing Imagery

2026-04-20 · Geet Sethi, Panav Shah, Ashutosh Gandhe, Soumitra Darshan Nayak arxiv

Diffusion models have emerged as powerful tools for a wide range of vision tasks, including text-guided image generation and editing. In this work, we explore their potential for object grounding in remote sensing imager…

Object LocalizationImage Generation

RECOUNT: Reference-guided Counting with Synthetic Visual Exemplars

2026-08-20 · Adriano D'Alessandro, Ali Mahdavi-Amiri, Ghassan Hamarneh arxiv

Text-guided zero-shot object counters excel at spatial localization but categorize poorly on novel or fine-grained classes: natural language is too coarse to fully specify visual identity, so they fail to separate visual…

SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding

2025-08-28 · Jiawen Lin, Shiran Bian, Yihang Zhu, Wenbin Tan 외 arxiv

3D Visual Grounding (3DVG) aims to localize objects in 3D scenes using natural language descriptions. Although supervised methods achieve higher accuracy in constrained settings, zero-shot 3DVG holds greater promise for …

3D Semantic SegmentationVisual Grounding