paper-with-me

홈 › Papers

A Pooling Approach to Modelling Spatial Relations for Image Retrieval and Annotation

2014-11-19 · Mateusz Malinowski, Mario Fritz

Over the last two decades we have witnessed strong progress on modeling visual object classes, scenes and attributes that have significantly contributed to automated image understanding. On the other hand, surprisingly little progress has been made on incorporating a spatial representation and reasoning in the inference process. In this work, we propose a pooling interpretation of spatial relations and show how it improves image retrieval and annotations tasks involving spatial language. Due to the complexity of the spatial language, we argue for a learning-based approach that acquires a representation of spatial relations by learning parameters of the pooling operator. We show improvements on previous work on two datasets and two different tasks as well as provide additional insights on a new dataset with an explicit focus on spatial relations.

📄 PDF Abstract BibTeX arXiv:1411.5190

Code (0)

등록된 구현이 없습니다.

Tasks

Image RetrievalRetrieval

Similar Papers 제목 키워드 기반

A CLIP-Hitchhiker's Guide to Long Video Retrieval

2022-05-17 · Max Bain, Arsha Nagrani, Gül Varol, Andrew Zisserman

Our goal in this paper is the adaptation of image-text models for long video retrieval. Recent works have demonstrated state-of-the-art performance in video retrieval by adopting CLIP, effectively hitchhiking on the imag…

RetrievalVideo RetrievalZero-Shot Action Recognition

All the attention you need: Global-local, spatial-channel attention for image retrieval

2021-07-16 · Chull Hwan Song, Hye Joo Han, Yannis Avrithis

We address representation learning for large-scale instance-level image retrieval. Apart from backbone, training pipelines and loss functions, popular approaches have focused on different spatial pooling and attention me…

AllImage RetrievalRepresentation LearningRetrieval

Compositional Sketch Search

2021-06-15 · Alexander Black, Tu Bui, Long Mai, Hailin Jin 외

We present an algorithm for searching image collections using free-hand sketches that describe the appearance and relative positions of multiple objects. Sketch based image retrieval (SBIR) methods predominantly match qu…

Image RetrievalPositionQuantizationRetrieval+2

BiC-Net: Learning Efficient Spatio-Temporal Relation for Text-Video Retrieval

2021-10-29 · Ning Han, Jingjing Chen, Chuhao Shi, Yawen Zeng 외

The task of text-video retrieval aims to understand the correspondence between language and vision, has gained increasing attention in recent years. Previous studies either adopt off-the-shelf 2D/3D-CNN and then use aver…

Cross-Modal RetrievalRelationRetrievalVideo Retrieval+1

Enhancing medical vision-language contrastive learning via inter-matching relation modelling

2024-01-19 · Mingjian Li, Mingyuan Meng, Michael Fulham, David Dagan Feng 외

Medical image representations can be learned through medical vision-language contrastive learning (mVLCL) where medical imaging reports are used as weak supervision through image-text alignment. These learned image repre…

Contrastive LearningCross-Modal RetrievalRelationRepresentation Learning+2