paper-with-me

Papers

Beyond Instance-Level Image Retrieval: Leveraging Captions to Learn a Global Visual Representation for Semantic Retrieval

2017-07-01 · CVPR 2017 7 · Albert Gordo, Diane Larlus

Querying with an example image is a simple and intuitive interface to retrieve information from a visual database. Most of the research in image retrieval has focused on the task of instance-level image retrieval, where the goal is to retrieve images that contain the same object instance as the query image. In this work we move beyond instance-level retrieval and consider the task of semantic image retrieval in complex scenes, where the goal is to retrieve images that share the same semantics as the query image. We show that, despite its subjective nature, the task of semantically ranking visual scenes is consistently implemented across a pool of human annotators. We also show that a similarity based on human-annotated region-level captions is highly correlated with the human ranking and constitutes a good computable surrogate. Following this observation, we learn a visual embedding of the images where the similarity in the visual space is correlated with their semantic similarity surrogate. We further extend our model to learn a joint embedding of visual and textual cues that allows one to query the database using a text modifier in addition to the query image, adapting the results to the modifier. Finally, our model can ground the ranking decisions by showing regions that contributed the most to the similarity between pairs of images, providing a visual explanation of the similarity.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Image RetrievalRetrievalSemantic RetrievalSemantic SimilaritySemantic Textual Similarity

Similar Papers 제목 키워드 기반

Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval

2026-04-07 · Yuxin Yang, Yinan Zhou, Yuxin Chen, Ziqi Zhang 외 arxiv

Composed Image Retrieval (CIR) has demonstrated significant potential by enabling flexible multimodal queries that combine a reference image and modification text. However, CIR inherently prioritizes semantic matching, s…

Image Retrieval

Beyond Pixels: A Training-Free, Text-to-Text Framework for Remote Sensing Image Retrieval

2025-12-11 · J. Xiao, Y. Guo, X. Zi, K. Thiyagarajan 외 arxiv

Semantic retrieval of remote sensing (RS) images is a critical task fundamentally challenged by the \textquote{semantic gap}, the discrepancy between a model's low-level visual features and high-level human concepts. Whi…

Cross-Modal RetrievalSemantic RetrievalImage Retrieval

Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization from Unposed Views

2026-06-28 · Mijin Yoo, In Cho, Subin Jeon, Jiwoo Lee 외 hf

A 3D scene is understood through its objects, not the primitives that compose them. Yet feed-forward reconstruction methods output dense, unstructured sets of points or Gaussians, leaving object-level structure to be rec…

Instance SegmentationNovel View Synthesis

Instance-level Image Retrieval using Reranking Transformers

2021-03-22 · ICCV 2021 10 · Fuwen Tan, Jiangbo Yuan, Vicente Ordonez

Instance-level image retrieval is the task of searching in a large database for images that match an object in a query image. To address this task, systems usually rely on a retrieval step that uses global image descript…

Image RetrievalRerankingRetrieval

Learning to Learn from Web Data through Deep Semantic Embeddings

2018-08-20 · Raul Gomez, Lluis Gomez, Jaume Gibert, Dimosthenis Karatzas

In this paper we propose to learn a multimodal image and text embedding from Web and Social Media data, aiming to leverage the semantic knowledge learnt in the text domain and transfer it to a visual model for semantic i…

Image RetrievalRetrieval