paper-with-me

홈 › Papers

Learning Spatial-context-aware Global Visual Feature Representation for Instance Image Retrieval

2023-01-01 · ICCV 2023 1 · Zhongyan Zhang, Lei Wang, Luping Zhou, Piotr Koniusz

In instance image retrieval, considering local spatial information within an image has proven effective to boost retrieval performance, as demonstrated by local visual descriptor based geometric verification. Nevertheless, it will be highly valuable to make ordinary global image representations spatial-context-aware because global representation based image retrieval is appealing thanks to its algorithmic simplicity, low memory cost, and being friendly to sophisticated data structures. To this end, we propose a novel feature learning framework for instance image retrieval, which embeds local spatial context information into the learned global feature representations. Specifically, in parallel to the visual feature branch in a CNN backbone, we design a spatial context branch that consists of two modules called online token learning and distance encoding. For each local descriptor learned in CNN, the former module is used to indicate the types of its surrounding descriptors, while their spatial distribution information is captured by the latter module. After that, the visual feature branch and the spatial context branch are fused to produce a single global feature representation per image. As experimentally demonstrated, with the spatial-context-aware characteristic, we can well improve the performance of global representation based image retrieval while maintaining all of its appealing properties. Our code is available at https://github.com/Zy-Zhang/SpCa

📄 PDF Abstract BibTeX

Code (1)

zy-zhang/spca 공식 구현 pytorch

Tasks

Image RetrievalRetrieval

Similar Papers 제목 키워드 기반

IAUnet: Global Context-Aware Feature Learning for Person Re-Identification

2020-09-02 · Ruibing Hou, Bingpeng Ma, Hong Chang, Xinqian Gu 외

Person re-identification (reID) by CNNs based networks has achieved favorable performance in recent years. However, most of existing CNNs based methods do not take full advantage of spatial-temporal context modeling. In …

Object CategorizationPerson Re-Identification

SARL: Spatially-Aware Self-Supervised Representation Learning for Visuo-Tactile Perception

2025-12-01 · Gurmeher Khurana, Lan Wei, Dandan Zhang arxiv

Contact-rich robotic manipulation requires representations that encode local geometry. Vision provides global context but lacks direct measurements of properties such as texture and hardness, whereas touch supplies these…

Self-Supervised LearningRepresentation Learning

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech

2024-12-16 · Rui Liu, Shuwei He, Yifan Hu, Haizhou Li

Visual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize the reverberant speech for the spoken content. The challenge of this task lies in understanding the spatial environment from t…

text-to-speechText to Speech

Dense Global Context Aware RCNN for Object Detection

2021-01-01 · Wenchao Zhang, Haoyu Xie, Mai Zhu, Chong Fu

RoIPool/RoIAlign is an indispensable process for the typical two-stage object detection algorithm, it is used to rescale the object proposal cropped from the feature pyramid to generate a fixed size feature map. However,…

Objectobject-detectionObject Detection

MTLDesc: Looking Wider to Describe Better

2022-03-14 · Changwei Wang, Rongtao Xu, Yuyang Zhang, Shibiao Xu 외

Limited by the locality of convolutional neural networks, most existing local features description methods only learn local descriptors with local information and lack awareness of global and surrounding spatial context.…

Indoor LocalizationTriplet